MCP and Assistant Tool Listings: Built to Get Selected

Yarrow WorksVerified agencyNew0 orders on this service
Emerging and Niche Channels · LLM plugin and tool-store optimization

We measure how often models actually select your tool, then rewrite the name, description, schema and error strings until that number moves.

About this service

Most listed tools are never chosen. On the first evaluation we run for a new client — 400 buyer-phrased prompts replayed against ChatGPT, Claude and Gemini — a live connector is selected on 8 to 15 percent of the prompts it was built to serve. The rest go to the model's own memory, to a competitor's tool with a plainer name, or to yours with the parameters filled wrong. That selection rate is the number this engagement moves, and it moves because of the manifest, not the marketing site. Four surfaces decide selection: The tool name, where a collision with an incumbent costs more than any listing copy can win back. The description field, which the model reads as an instruction rather than a blurb: it needs a trigger condition and a negative case, in the form of "use this when the user names a clinic, a policy number or an order ID; do not use it for general symptom questions". The JSON Schema, where a free-text parameter with a paragraph of explanation is the largest single source of malformed calls and an eight-value enum usually fixes it outright. And the error strings your server returns, because the model reads them and decides from the wording whether to retry, apologise to the user, or abandon you mid-turn. A 500 with an empty body ends the conversation. A 422 naming the field that was wrong gets a corrected second call about four times in five. Latency belongs in the same list. We measure p95 tool latency under concurrent load and treat anything past two seconds as a selection problem rather than an engineering one, because a model that has timed out on you once in a session routes around you for the rest of it. The evaluation harness: The prompt set is written from your support tickets and site search logs, in the customer's own words, not from a keyword tool. We score three things separately — was your tool selected, was the right tool selected, and were the parameters filled correctly — because they have different fixes and averaging them hides which one is broken. The harness ships as a repo in your GitHub organisation: prompt set, runner, rubric. You keep it. We do not host it and there is nothing to keep paying us for once it runs. Re-running matters more than the rewrite. Routing changes without announcement, and a description that won selection in March can decay by June with nothing changed on your side. Monthly is the slowest cadence we would defend. Listings are a separate problem: Directory review is not selection. Submissions get rejected for privacy policies that do not cover the data the tool actually touches, retention language that contradicts the schema, and screenshots showing a flow the tool does not perform. We prepare the submission and handle one resubmission. What we do not do: We do not build or host your server, and we do not take write access to it. We specify the changes and your engineers land them, which is also why the loop only works if someone on your side can ship a schema change inside a fortnight. We do not write descriptions designed to pull the model into queries the tool has no business answering. That tactic works for a few weeks, ends in delisting, and we will not take the rebuild afterwards. Who this is not for: Products without a working authenticated endpoint — there is nothing to evaluate. Teams on a quarterly release train. And anyone whose real goal is to be mentioned in AI answers generally: that is a different discipline, sold separately on this profile, and treating the two as one thing costs a quarter. We work in English and Mandarin, and run the same harness against Qwen and Doubao when the client sells into China. Scoring is identical; the prompt set is written natively rather than translated.

Scope

Target market
Worldwide, Southeast Asia, Singapore
Working language
English, Chinese (Simplified)
Industry
B2B SaaS, Developer tools, Marketplaces, Health and wellness
Engagement model
One-off project
Turnaround
1 month or more
Seller type
Full-service agency

What the seller needs from you

  1. 1Link to your live manifest, OpenAPI spec or MCP schema.
  2. 2Export of the last 90 days of support tickets or site search queries.
  3. 3Which tools do you believe you lose selection to?
  4. 4Who on your side lands schema changes, and how often do they ship?

Asked at checkout. Delivery time starts once you answer, not when you pay.

Reviews

No reviews on this service yet.

Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.

Other sellers offering llm plugin and tool-store optimization

See all →

Starting at $7,000