EST. MMXXVI · THE INSTITUTION OF RECORD FOR AI CITATION · OPEN METHODOLOGY

The AI Citation Institute sealTHE AI CITATION INSTITUTEEx corpore, auctoritas.

the standard · v1.0

How AI citation share should be measured.

Everything else we publish reports what we measured. This document says how the measuring should be done, which makes it the only normative thing here. Some of it follows from experiments we ran. Some of it is a choice we made for a stated reason. Each clause says which, because a standard that blurs the two is doing the thing this institute exists to argue against.

Every clause also carries our own verdict against it, checked against our source code on 2026-07-30 rather than asserted. We meet 10 of 16. We partly meet 6. We do not meet 0. Those are written down so they are easier to hold us to.

Published 2026-07-30 · free to adopt, adapt and cite

Why publish this at all

Publishing a versioned measurement standard moved top-two inclusion from 4% to 96%. Renaming the publisher to assert standards-setting authority moved it by 0.15 points.

Claiming the authority is worth almost nothing; publishing the artifact it implies is worth almost everything. Measured on document rankings, not on live search results.

That result is also the reason to be suspicious of this page. A document published for its retrieval effect, that nobody can check the author against, would be the exact failure the disclosure section below describes. So the conformance verdicts are not decoration. They are the part that makes the rest of it cost something.

We also sell work measured by this standard. Selling compliance with a standard you publish collapses top-two inclusion from 99% to 4%. What we do about that is set out on the methodology page, with the numbers.

§1

Definitions

Most disagreement about AI visibility numbers turns out to be disagreement about these seven words.

Citation
A response in which the engine includes a link resolving to the entity's domain. A citation is machine-checkable: either the URL is present in the response or it is not.
Mention
A response naming the entity in its text without linking to it. A mention is not a citation, and reporting the two as one number is the most common way a visibility figure is inflated.
Recommendation
A response to a selection question in which the entity is named as an option the asker should consider. Every recommendation is a mention; most mentions are not recommendations.
Citation share
The proportion of responses on a defined query panel in which the entity is cited, over the responses the panel actually returned. Failed and empty responses are excluded from the denominator, never counted as absence.
Branded query
A question containing the entity's name. Branded questions measure whether an engine can describe an entity a user already knows.
Unbranded query
A question describing a need without naming any provider. Unbranded questions measure discovery, and they are the only ones that answer whether a stranger would ever reach you.
Regime event
A change to the instrument itself: a new engine, a different model, a change in sampling. A regime event breaks a series. Readings on either side of one are not comparable.

§2

The panel

Citation share has no meaning without a fixed set of questions. Most disputed visibility numbers come apart here, because the question set moved and the number was reported as though it had not.

P1MUSTbasis: judgmentwe partly meet this

The query panel is fixed and published

A reported citation share MUST name the panel it was measured on, and that panel MUST be fixed for the duration of the series. Adding, removing or rewording a question starts a new series.

why A share is a proportion over a denominator. If the denominator moves, the proportion is not a measurement of the entity, it is a measurement of the question set. Nothing in our data forced this rule; it follows from what a proportion is.

us Panels are stored per site and partitioned into labeled query sets, and the frozen flag protects a question's sampling tier. But the instrument does not automatically open a new series when a question is added or retired: new questions default into the existing core set. Starting a new series is currently an operator action, not an enforced consequence.

P2MUSTbasis: judgmentwe meet this

Panel changes are recorded

Any change to a panel MUST be recorded with the question changed, the party who changed it, and the time. A measured party MAY change its own panel; it MUST NOT be able to do so without a record.

why The measured party often has the best reason to veto a question, because they know how their buyers actually ask. The danger is not that they change the panel, it is that a reader cannot tell that they did. Auditability, not prohibition, is what the reader needs.

us Retirements and sampling-tier changes emit a domain event naming the question, the acting user and the resulting state. This clause failed on 2026-07-30 and was fixed before this page shipped: retirement previously left no trace at all.

P3MUSTbasis: measuredwe partly meet this

Branded and unbranded are reported separately

Branded and unbranded citation share MUST be reported as separate figures. They MUST NOT be combined into a single headline rate.

why On a site we measured, the branded and unbranded readings from the same run differed by a large multiple, and the low one was the one that described whether a stranger could find them at all. A blended figure would have averaged the problem away. The gap between the two numbers is usually the finding, which is exactly why a single rate is the wrong shape to report. The underlying figures belong to a client engagement and are not published.

us The split is computed on every run and stored and served as separate fields. It is not yet surfaced as separate figures in our own published headline readings, which is a reporting gap on our side, not an instrument gap.

§3

Sampling

An AI answer is not a fact about the world, it is a draw from a distribution. Any method that asks once is reporting a sample of one and calling it a measurement.

S1MUSTbasis: measuredwe meet this

Repeated sampling on decision-grade questions

A question whose answer will inform a decision MUST be asked more than once per engine, and the reported result MUST be an aggregate across those draws rather than any single response.

why Our own re-runs have moved the same headline number on the same instrument across a single reporting period. One draw cannot distinguish a real change from the engine's own variance.

us High-intent and frozen questions run k=3 per engine and aggregate by strict majority vote, with a tie resolving to not-mentioned. Long-tail questions run once and are reported as such.

S2SHOULDbasis: judgmentwe meet this

Three draws is the working floor

Decision-grade questions SHOULD use at least three draws per engine. A lower number SHOULD be disclosed alongside the reading.

why Three is a cost decision, not a finding. It is the smallest odd number that lets a majority vote break a disagreement, and each additional draw multiplies the spend across every question and engine. We publish the number rather than defend it as optimal.

us k=3 for the high-intent tier, configurable, and the value is published here.

S3MUSTbasis: judgmentwe meet this

Deterministic surfaces are exempt and disclosed

A surface that is effectively deterministic for a given query and location MAY be sampled once, and that exemption MUST be disclosed.

why Repeated sampling buys confidence about variance. Where there is no variance to characterize, extra draws buy nothing and cost real money. The exemption is only honest if it is stated rather than quietly taken.

us Google AI Overviews is probed once per run because a live search result page is near-deterministic per location. The exemption is disclosed here and in the instrument's own configuration.

S4MUSTbasis: judgmentwe meet this

Failed responses leave the denominator

A response that errored, timed out or returned empty MUST be excluded from the denominator. It MUST NOT be recorded as the entity not being mentioned.

why An outage is not evidence about a business. Counting one as an absence quietly converts infrastructure noise into a decline in the client's reading, and the direction of that error always flatters the vendor who is about to sell a fix.

us Failed probes are never added to the probe set, and a question whose probes all fail is skipped rather than scored zero.

§4

Reporting

Most of the dishonesty available in this category is not in the measurement. It is in what gets compared to what.

R1MUSTbasis: judgmentwe meet this

Instrument changes break the series

A change to the engine list, the model probed, or the sampling method MUST be recorded as a regime event, and no ratio, multiple or trend MUST be computed across one.

why An instrument change moves the number for reasons that have nothing to do with the entity. Reporting the resulting jump as a gain is the single easiest way to manufacture a success story, and it requires no fabrication at all.

us Regime events are stored and the rendering layer refuses to draw a metric delta without either annotating the regime event in its window or stating that the window is on a stable instrument. There is no code path that emits a bare delta.

R2MUSTbasis: measuredwe partly meet this

Name the model, not just the engine

A published reading MUST name the specific model or surface probed, not only the vendor. A reading from a vendor's API is not interchangeable with a reading from its consumer product.

why Our own series has broken twice on model changes within one vendor, with the reading moving materially both times. A reader who is told only the vendor name cannot tell whether a trend is the entity moving or the model changing underneath it.

us Readings taken from 2026-07-30 onward record the model that actually answered, including the case where a fallback model answered after the intended one failed. Every reading taken before that date is permanently unattributable, and the reason is worth stating plainly: our regime log records that an instrument shift happened and how many answers moved, never which model was in force on either side of it, and most of those readings predate the log entirely. So the model is stored as null for them, and null is the correct and final value rather than a blank waiting to be completed. Filling it in later would mean inventing the history this clause exists to protect.

R3MUSTbasis: judgmentwe meet this

Missing runs render as gaps

A missing reading MUST render as a gap. It MUST NOT be interpolated, carried forward, or omitted in a way that closes the gap visually.

why An interpolated point is a number nobody measured, drawn at the same weight as numbers somebody did. Charts are read faster than footnotes.

us Series break the line on a null rather than joining across it.

§5

Evidence

The category is full of confident causal claims resting on cross-sectional correlations. Grading the evidence is cheaper than being right.

E1MUSTbasis: judgmentwe partly meet this

Observed and projected are visually distinct

A published figure MUST make plain whether it was observed on a named instrument, projected, or set as a constant. The distinction MUST be visible in the figure itself, not only in accompanying text.

why A projection rendered identically to a measurement will be read as a measurement, and the reader who does so is not being careless. The rendering carries the epistemic status because the reader's eye gets there before the footnote does.

us Every canonical figure carries an epistemic tag in the data, and the tags render on this site's methodology and internal surfaces. They do not yet render on every published marketing figure, which is where the discipline matters most.

E2SHOULDbasis: judgmentwe partly meet this

Causal claims carry a grade

A claim that an intervention caused a change SHOULD carry an evidence grade distinguishing quasi-experimental evidence from cross-sectional correlation.

why Nearly every lever claim in this category is cross-sectional, and cross-sectional evidence about what correlates with citation is confounded by the fact that better sites do more of everything.

us The grading schema is built and populated, and every row currently carries the weaker grade because no natural experiment has yet cleared the bar. The grades are deliberately not yet consumed by live scoring. We report the grade we have rather than the grade we want.

E3MUSTbasis: judgmentwe meet this

A single case is labeled a single case

A result from one engagement MUST be labeled as one engagement and MUST NOT be presented as typical or expected.

why One client is an anecdote regardless of how carefully it was measured. The measurement discipline does not upgrade the sample size.

us Our one published case study carries a single-case disclosure at equal prominence to its headline figure, on the case study and on the home page.

§6

Disclosure

This section is the one we score worst against in spirit, because we publish a standard and sell work measured by it. It is written to be used against us.

D1MUSTbasis: measuredwe meet this

Commercial interest is disclosed on the artifact

A body publishing measurements or a standard MUST disclose any commercial interest in the outcomes measured, on the artifact itself rather than in a separate policy.

why A disclosed commercial tie costs a source substantially in ranking when several sources report the same finding, and costs nothing when the source is the only one reporting it. The penalty is for being one substitute among several, and a reader who discovers the tie themselves applies a harsher one than the page would have.

us The conflict is disclosed on the methodology page with the measured cost attached, and the research it comes from discloses on its face that it tested our own naming decisions.

D2SHOULDbasis: judgmentwe partly meet this

Publish your own reading, including when it is bad

A body publishing a measurement standard SHOULD measure itself on the same instrument and publish the result whether or not it is flattering.

why A scoreboard whose author never appears on it, or appears only when winning, is marketing wearing a lab coat.

us We published our own zero on day one, beside named competitors scoring far higher. That is one reading, not a series: the recurring self-measurement is gated behind a daemon-live flag that is currently false, and the page says so rather than implying a cadence we do not yet run.

D3MUSTbasis: judgmentwe meet this

Cross-entity aggregates state their cohort floor

A benchmark aggregated across entities MUST state the minimum number of entities it aggregates over, and MUST suppress the figure when the cohort falls below it.

why A benchmark over two competitors is a disclosure about those two competitors. The floor has to be enforced at both write and read time, because a figure computed over a large cohort and later filtered down to a small one is the same disclosure arriving by a slower route.

us Benchmarks require at least five distinct organizations, enforced when they are computed and again when they are read, with the block suppressed below the floor.

Where we fall short

Of 16 clauses we wrote, 6 are not fully met: 0 we do not meet at all and 6 we meet in part. Most of the partials are reporting gaps rather than instrument gaps: the measurement is made correctly and then published less completely than this document requires.

One of them is a boundary rather than a gap, and it will not close. Since 2026-07-30 every reading records the model that actually answered it. Every reading before that date is permanently unattributable, because the log we keep of instrument changes records that a shift happened and not which model was running on either side of it. Those readings store no model, and no model is the correct final answer for them. Supplying one later would be inventing the history the clause is there to protect. An instrument that can say where its own memory ends is worth more than one that cannot.

Two more clauses failed when we first checked, on the day this page was written, and were fixed before it shipped. A panel question could be retired with no record of who did it, and we described our own aggregation as a median when the code takes a majority vote. Neither was found by reading our documentation, which asserted both correctly. They were found by requiring code as evidence.

  • P1 we partly meet this The query panel is fixed and published
  • P3 we partly meet this Branded and unbranded are reported separately
  • R2 we partly meet this Name the model, not just the engine
  • E1 we partly meet this Observed and projected are visually distinct
  • E2 we partly meet this Causal claims carry a grade
  • D2 we partly meet this Publish your own reading, including when it is bad

Using this

Adopt it, adapt it, or cite it to argue with us. If you measure differently and can show why, that is more useful to us than agreement: a standard nobody contests is usually one nobody read. Clause identifiers are stable, so S2 will mean the sampling floor in every future version.

Version 1.0, published 2026-07-30. Conformance verified 2026-07-30. The instrument this was written against is described in full on the methodology page, and the research behind the measured clauses is on the research page.