The probe, step by step
Measuring AI recommendations is an instrument problem. One query on one engine on one day tells you almost nothing, because assistants sample their own answers and drift. A probe fixes the instrument so the number moves only when reality moves.
| Step | What you do | Why it matters |
|---|---|---|
| 1. Build the query set | Write the real questions buyers ask an assistant — not your keywords, the phrasing a customer would actually type. | You measure the questions that decide purchases, not the ones that flatter you. |
| 2. Run the panel | Send each query across the frozen five-engine panel — ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews — several times per engine. | Answers vary run to run; sampling several times per engine turns one noisy shot into a stable read. |
| 3. Score the answer | Read each response and label it mentioned, recommended, or absent. | Being named and being recommended are different outcomes and have to be counted apart. |
| 4. Track weekly medians | Roll the scores into a share and plot it as weekly medians, annotating any change to the instrument. | Medians cancel jitter; annotations keep an instrument change from reading as a real shift. |
The panel is frozen on purpose — same engines, same sampling — so week-over-week numbers are comparable.
Why one query is noise
Assistants don't return a fixed answer. Ask ChatGPT the same question twice and the wording, the brands named, and the order can all change — that's sampling, built into how these models generate text. So a single query is a coin flip, and a single-engine check misses that ChatGPT, Claude, Gemini, Perplexity, and AI Overviews each have their own retrieval and their own biases.
The fix is boring and it works: sample each query several times per engine, run every engine in the panel, and read the result as a weekly median. That's why The AI Citation Institute measures across all five engines every week rather than spot-checking one. A number you can't reproduce isn't a measurement.
What to capture per answer
For every sampled answer, record enough to reproduce the read later and to separate a mention from a recommendation. Store the fields below alongside the raw response text.
| Field | What it records |
|---|---|
| Query | The exact buyer question that was sent. |
| Engine | Which of the five panel engines produced the answer. |
| Sample | The run index, since each query is sampled several times per engine. |
| Verdict | Mentioned, recommended, or absent. |
| Competitors named | Every rival brand the answer surfaced, for share-of-voice. |
| Timestamp | When the sample ran, so it rolls into the right weekly median. |
The full instrument — query panel, sampling cadence, and scoring rubric — is documented at /methodology.
Mention rate vs recommendation rate
Keep the two numbers separate. Mention rate is how often the assistant names you at all; recommendation rate is how often it actively steers a buyer toward you. They move differently, and the gap between them is usually the real story.
For Gavelist, one ongoing The AI Citation Institute client, the probe showed a 58.5% mention rate against the nearest rival's 38.8% — a clear lead on being named. Active recommendation ran about 13%. Getting mentioned is common; getting recommended is scarce, and that's the number worth closing.
Common questions
Why sample each query several times per engine instead of once?+
Assistants sample their own output, so the same question returns different answers run to run. One shot is noise. Sampling several times per engine and reading weekly medians gives you a number you can reproduce and trust.
Which engines should the panel cover?+
A frozen five-engine panel: ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews. Each has its own retrieval and biases, so checking one engine misses most of the picture. Keeping the panel frozen keeps week-over-week numbers comparable.
What's the difference between a mention and a recommendation?+
A mention is the assistant naming your brand at all. A recommendation is the assistant actively pointing a buyer to you. They're scored separately because they move differently — for one The AI Citation Institute client, mention rate hit 58.5% while active recommendation ran about 13%.
keep reading
- Our audit of the pages AI actually cites→
- How much does AI visibility software cost?→
- Our nine-business study: AI knows who you are, it will not volunteer you→
- The full measurement instrument→
- Live AI visibility scoreboard→
- What is AI citation tracking?→
- What is AI citation share?→
- How to measure AI visibility→
- What does an AI visibility audit include?→
See what AI says about your brand right now — free, on a live engine.
Scan your domain