What the standard covers
Version 1.0 is organized into five parts: definitions, then clauses on the panel, sampling, reporting, evidence and disclosure. The definitions do more work than they look like they do, because most arguments about AI visibility numbers turn out to be arguments about whether a mention counts as a citation, or about what the denominator was.
| Group | What it governs | The clause most often broken elsewhere |
|---|---|---|
| Panel | Which questions are asked, and what happens when that set changes | Branded and unbranded results reported separately, never blended into one rate |
| Sampling | How many times each question is asked, and how draws are combined | Failed responses excluded from the denominator rather than counted as an absence |
| Reporting | What may be compared to what | No trend computed across a change to the instrument itself |
| Evidence | How confident a published claim is allowed to sound | Projections rendered visibly differently from measurements |
| Disclosure | What the publisher must say about its own interests | Commercial interest disclosed on the artifact, not in a separate policy |
Clause identifiers are stable across versions, so a reference to S2 will always mean the sampling floor.
Why a measured claim and a judgment call are labeled differently
A standard is a normative document. It says how something should be done, which means some of it cannot come from measurement at all. The rule that failed responses leave the denominator follows from what a proportion is, not from an experiment. The rule that decision-grade questions get at least three draws is a cost decision, and three is simply the smallest odd number that lets a majority break a tie.
Other clauses do rest on findings. Reporting branded and unbranded separately is there because on a site we measured, the two readings from the same run differed by a large multiple, and the lower one was the one describing whether a stranger could find the business at all. A blended figure would have averaged that problem away. The underlying figures belong to a client engagement and are not published.
Every clause says which of the two it is. A standard that presents a preference as a finding is doing the thing measurement discipline exists to prevent.
The publisher grades itself, and does not pass
Each clause carries a verdict on whether The AI Citation Institute's own instrument meets it, checked against source code rather than asserted. Three clauses are not fully met, and the largest is that published readings name the engine but not the specific model behind it, so a reading from before a model change cannot be distinguished from one after it except by hand.
Two further clauses failed when the check was first run and were fixed before the standard was published: a question could be retired from a panel with no record of who removed it, and the aggregation across draws was described as a median when the code takes a majority vote. Neither was caught by reading the documentation, which described both correctly. They were caught by requiring code as evidence.
Publishing the gaps is the mechanism. A standard the author cannot be held to is a marketing document with clause numbers.
Common questions
Is the AI Citation Measurement Standard free to use?+
Yes. It is published to be adopted, adapted, or cited in disagreement. Contesting a clause with evidence is more useful to its authors than agreeing with it.
Does the standard require a specific tool?+
No. It constrains the method, not the vendor. Any instrument that fixes its panel, samples repeatedly, separates branded from unbranded, and breaks its series on instrument changes can conform.
What is a regime event?+
A change to the measuring instrument itself, such as a new engine, a different underlying model, or a change in sampling. Readings on either side of a regime event are not comparable, and the standard forbids computing a trend or a multiple across one.
Why must branded and unbranded be reported separately?+
Because they measure different things. Branded questions test whether an engine can describe a business someone already named. Unbranded questions test whether a stranger would ever reach it. On a site we measured the two readings differed by a large multiple in the same run, and a single blended number would have hidden the gap that mattered.
Does the publisher meet its own standard?+
Not fully, and the page says so clause by clause. Three of sixteen clauses are not currently met, including persisting the specific model behind each historical reading.
keep reading
See what AI says about your brand right now — free, on a live engine.
Scan your domain