A frozen prompt set run monthly across five assistants, reported with intervals so you can tell real movement from model noise.
About this service
Ask an assistant the same buying question thirty times and you will not get thirty identical answers. Across the accounts I measure, a brand that appears at all shows up in roughly 6 to 19 of 30 repeats of one prompt, run the same day, with nothing changed on your site. A screenshot of a good answer is therefore not evidence of anything. This programme turns that variance into a monthly figure with a stated interval, so your team can tell a real shift from a noisy Tuesday.
What runs every month:
A prompt set agreed with you in week one and then frozen. Freezing matters more than size, because a set that changes month to month cannot produce a trend. Each prompt is run thirty times per surface per market, in a fresh session, with browsing forced on and again with it off, since the two conditions answer differently and mixing them hides the cause. Default surfaces are ChatGPT through the OpenAI Responses API with the web search tool, Claude through the Anthropic Messages API with its web search tool, Gemini through Google AI Studio with grounding enabled, Perplexity through the Sonar API, and Google AI Overviews and AI Mode captured from live result pages, since no API exposes them.
How the numbers are built:
Every call lands as one row: prompt id, surface, model version string, browsing state, timestamp, full answer text, matched brand entities, cited URLs, and the sentence position of your first mention. Four readings come out of those rows. Mention rate per prompt cluster with a 95 percent interval. Citation share, which names the pages assistants actually quote when they mention you, and that is usually a review site or a forum thread rather than your own domain. First-mention position, because being named in the last line of an answer is not the same as being the recommendation. And a version log, so when a model version string changes and your numbers move the same week, the report says so instead of crediting your marketing.
The honest part about variance:
These systems are non-deterministic, and they are also moving under you. Sampling temperature, retrieval, rollout cohorts and account personalisation all shift without notice. Three consequences to accept before you buy. First, small changes are not detectable at this sample size: I will defend a swing of around fifteen points in mention rate, not three, and I will say plainly when a change sits inside the interval. Second, API answers are not app answers. Consumer apps carry their own system prompts, memory and account history. I measure API surfaces plus captured AI Overviews, and every chart is labelled with what it is. Third, a flat month is a result. I report no movement as no movement rather than dressing it up.
Not included:
No content writing, no digital PR, no outreach, and no work whose purpose is to move the numbers I report. I will not grade my own homework, and any measurement supplier who also sells the fix has a reason to find progress. Also outside scope: paid media reporting, keyword rank tracking, and general site audits.
Who this is not for:
Brands nobody asks assistants about yet. If your category questions do not name you in any answer across a baseline run, the honest advice is to spend the budget on being findable and come back in two quarters, and I will say that after the first cycle rather than sell you eleven more. It is also wrong for teams who need the line to rise this quarter. The first three cycles establish a baseline; decisions come after that.
Who does the work:
Me. Wei-Lin Tan, every run and every read, with no analyst pool between you and the data. Earlier this year I told a Singapore payments client to drop sentiment scoring from their programme: their scores tracked how politely each model writes, not what their market thinks of them, and paying monthly to watch that was waste. The same judgement gets applied to your set.