EST. MMXXVI · THE INSTITUTION OF RECORD FOR AI CITATION · OPEN METHODOLOGY

The AI Citation Institute sealTHE AI CITATION INSTITUTEEx corpore, auctoritas.

the instrument

We publish how we measure. The transparency is the product.

Most vendors in this category hand you a score and hide the ruler. We do the opposite: here is the exact instrument, the rubric that scores each page, and the rails that keep every number honest. If you can cite it back to us, it's on this page.

How the panel reads.

The panel asks your buyers' real questions the way your buyers ask them, on the engines your buyers actually use, enough times to tell signal from noise.

  1. 01

    A frozen query panel

    Each site's query panel is frozen so the same questions are asked every run; changing it starts a new labeled series.

    Change the panel and you start a new labeled series. The questions don't drift under the measurement.

  2. 02

    Five engines, every run

    Panel probes ChatGPT, Claude, Gemini, Perplexity, and Google AI answers.

    AI answers vary by user and over time. We report panel medians, not guarantees.

  3. 03

    k draws per query, median reported

    Panel runs each high-intent query k=3 times per engine (VISIBILITY_TIER1_K) and takes a strict majority vote across the probes; long-tail queries run once. google_aio is exempt from k-sampling because a live SERP is near-deterministic per location.

    One draw is noise. The reading is the majority verdict across k draws, and a tie resolves to not-mentioned, which is the conservative direction.

  4. 04

    Scored per query

    Each query on each engine resolves to one of three states — mentioned, recommended, or invisible — and rolls up to a composite panel score. We report the median, not the best single draw.

    One draw is noise. The reading is the median across the draws.

measured Engines don't behave alike. On our own runs, Gemini (Google-grounded) cited us most often, ChatGPT (Bing browse) has the weakest brand-authority gate and is the most reachable for a below-threshold brand, and Perplexity rewards citation density and title match.

Observed Gavelist rates at one point in time, not universal or guaranteed.

the rubric

Three layers, then a gate.

Every page we ship is scored against a declared target query on a 100-point rubric built from 180+ controlled queries across three unrelated industries, cross-checked against the published Princeton, Ahrefs, and Yext research. Retrieval gets you found; selection gets you quoted.

LayerNameWeightWhat it scores
GATEEligibilitypass / failIndexed by Google (Gemini's source) and Bing (ChatGPT's source), rendering clean, no placeholder markers. A page that fails the gate is flagged for repair before it is scored — you can't be selected from a result you were never retrieved into.
L1Entity & authority40 ptsThird-party mentions (the single strongest signal — per the Ahrefs 75,000-brand study, brand mentions correlate with AI visibility about 3× stronger than backlinks), author credentials, entity consistency, on-page social proof, non-promotional tone. Mostly off-page — the part no page edit fixes.
L2Retrieval25 ptsOn-page SEO scored against the AI-likely query: semantic title–query alignment, page type matching intent, and freshness with substantive change. This is where a page becomes retrievable.
L3Selection35 ptsWill the model quote you once it has retrieved you? Extractable numeric claims, a standalone-quotable passage, cited statistics, first-party data the model can't generate itself, structured tables. The genuinely novel GEO layer.

We report two numbers, never one: a citation-probability score (how likely the page is to be cited) and an on-page action-priority score (which controllable gaps to fix next). One number hides whether your bottleneck is fixable by editing a page or by building your entity off-page.

What we'll sell you — and what we won't.

confirmed levers

  • measured Inbound internal links are the strongest on-page lever we've measured: in our own per-post audit, posts with zero inbound internal links were self-cited 4.5% of the time, and posts with ten or more were self-cited around 44% of the time.

    Our own corpus, one point in time. A correlation on our data, not a guarantee for yours.

  • measured Semantic title–query alignment: in our cross-niche sample, cited pages had titles that were semantic equivalents of the target query and non-cited pages did not.

  • measured Freshness with substantive change: content updated within 30 days is cited more often than stale content — but a date-only bump does not count, and re-bumping updated_at in bulk can hurt.

    Substantive means at least one stat or section actually changed. Date-only bumps are worthless and bulk-bumping can be punished.

hygiene, not levers

  • FAQPage schema is hygiene, not a lever. It helps a machine parse a page it already retrieved; it does not make an invisible page get cited. We add it; we don't sell it as the reason you'll win.

  • An llms.txt file is hygiene, not a lever. There's no measured evidence that publishing one changes whether AI recommends you. We keep one tidy; we don't pretend it's the strategy.

We ship these because tidy is better than untidy. We refuse to bill them as the reason your answer changes, because on our own data they aren't.

the honesty rails

The disciplines that keep the number honest.

Measured vs modeled, in the pixels

A measured number is observed on the named instrument at a stated time — it renders solid and saturated. A projection or benchmark renders desaturated, dashed, and hollow. A forecast can never be mistaken for a fact because the rendering itself says which it is.

Regime annotation on every trend

When we change how we probe — a new engine, more draws, a different sampling — that is an instrument change, not a client win. Every such change is drawn as a first-class vertical line on the chart, and we never compute a multiple across one.

Nondeterminism is disclosed, not hidden

AI answers vary by user and over time. Our own re-runs have swung the same headline number on the same instrument. That's why we report panel medians across k draws, not a single lucky answer, and why every superlative is scoped to the published instrument.

Outages show as gaps

When an engine is down or a run is missing, the chart shows a gap — not an interpolated guess. We publish our own reading live and never retouch numbers.

The instrument is separated from the service

We publish a measurement standard and we also sell work that helps sites do better against it. That is a conflict, we measured what it costs, and the separation below is our answer to it.

the conflict of interest

We publish the standard and we sell the work. Here is what that costs.

A body that publishes a measurement standard and also sells help passing it has a problem, and we have it. We went and measured how large it is, on ourselves.

measured Selling compliance with a standard you publish collapses top-two inclusion from 99% to 4%.

measured An explicit editorial firewall recovers 78% of that collapse (95% CI 70 to 87).

measured A scarcity pitch, one client per vertical, recovers 8% of the same collapse.

measured A disclosed commercial tie costs 2.70 ranking points when several publishers report the same finding, and costs nothing when the publisher is the only source for it.

Tier 2 evidence. Holds pooled across model families but is NOT statistically significant on GPT alone, and it measures the WORDING of a disclosure, not the existence of an organizational separation. Language moved the ranking; do not read it as proof a firewall works.

So a firewall statement is not the answer. It is a sentence, and our own data says sentences buy less than they appear to. These are the structural commitments instead, each of which you can check.

The panel cannot be changed silently

You can retire a question that does not reflect how your buyers really ask, and you should. What you cannot do is remove one without a trace: every retirement and every change to a question's sampling tier is written to the event log with who made it and when. Ours are in the same log as yours.

We publish our own zero

The day we onboarded ourselves we scored 0.0% on our own panel, beside named competitors scoring far higher, and we published it. A scoreboard you can only win on is not a scoreboard.

Buying the service does not buy a number

Optimization work changes what a site says and how it is corroborated elsewhere. It has no path into the measurement: the panel, the sampling and the scoring run the same way whether or not you are a client.

The research behind this is self-interested, and says so

The study these figures come from tested our own naming and brand-architecture decisions. We disclose that on the paper and here, because a commercial tie costs a source far more when a reader discovers it than when the source states it.

Source: The AI Citation Institute, 'What makes an llms.txt get cited', §5; independently recomputed in the fresh-eyes adversarial review, 2026-07-27 (n=54 rankings, pooled).

what actually gets fetched

Nothing fetched the llms.txt file.

measured Across 18 days and roughly 15,000 requests from AI-engine crawlers and answer-time agents to a live production site, not one requested /llms.txt.

One site, one window. What the answer-time agents did fetch was the homepage, one statistics-dense blog post, and individual product pages. Serve the file, but put load-bearing facts in rendered HTML.

Source: Production access logs, 2026-07-09 to 2026-07-27: GPTBot 436, OAI-SearchBot 577, ChatGPT-User 1,406, ClaudeBot 747, Claude-User 197, PerplexityBot 690, Googlebot 10,007, GoogleOther 773, Google-Extended 7. The only non-human fetchers of the file were bingbot (3), an SEO crawler, and our own monitoring.

The rails, applied.

Here is the discipline on a real chart. This is our own frozen five-engine panel for Gavelist. The two dashed vertical lines are instrument changes, drawn as first-class marks; the line never states a clean multiple across them.

gavelist · own five-engine panel · weekly medians · ongoingmeasured
04-28 · instrument06-15 · instrumentApr 20Jul 20

Solid = measured on one instrument. Dashed segments cross a measurement-mode change (04-28 and 06-15): how we probe changed, not the site. We don't compute a multiple across them.

The value of this series is its shape and its honesty, not any single figure. A reader can see the climb and see exactly which parts of it were the site improving versus the ruler changing.

One engagement, started 2026-04-23 and still running. A single case — not typical or guaranteed.

read the full case study →

Point the same instrument at your brand.