The four things a real measurement needs
A vanity 'AI score' from a single prompt is worthless because the same prompt returns a different answer on the next run. A defensible measurement fixes four variables so movement means something.
| Ingredient | What it means | Why it matters |
|---|---|---|
| Frozen panel | The same buyer questions every run | Comparability — the ruler can't drift under you |
| Multi-engine | Every engine your buyers actually use | One engine is a keyhole; buyers use several |
| k-sampling | Each query asked several times per engine | Damps single-draw nondeterminism; median = signal |
| Citation capture | Which URLs the engine cited | Lets you attribute a move to a specific page |
Three states, not one number
Score each answer as recommended, mentioned, or invisible. 'Recommended' and 'mentioned' are different outcomes — being named in a list is not the same as being the pick — and collapsing them into a single score hides the gap you most need to close.
Roll the per-query states up to a composite panel score, and always report the median across draws rather than the best single answer you happened to catch.
Annotate every change to the instrument
The moment you add an engine, change how many times you sample, or edit the panel, you have changed the ruler — not the brand. If you don't mark that on the chart, a measurement artifact looks like a win.
Draw instrument changes as first-class vertical lines and never compute a ratio across one. This is the discipline that separates an honest trend line from a flattering one.
AI answers vary by user and over time. A panel reports medians, not guarantees.
Common questions
Can't I just ask ChatGPT once and screenshot it?+
You can, but it won't be reproducible. Ask the same question again and the wording — and sometimes whether you're mentioned at all — changes. A single screenshot is an anecdote; a frozen, k-sampled panel across engines is a measurement.
How many engines do I need to track?+
The ones your buyers use. the AI Citation Institute's panel probes five: ChatGPT, Claude, Gemini, Perplexity, and Google AI. They don't behave alike — Google-grounded engines and Bing-grounded engines cite different sources — so a single engine gives a misleadingly narrow read.
How often should I measure?+
Weekly is a sensible cadence for a frozen panel: frequent enough to catch movement, infrequent enough that noise averages out. What matters more than frequency is that the panel stays frozen between runs.
keep reading
See what AI says about your brand right now — free, on a live engine.
Scan your domain