What Actually Makes a Source Get Cited by an AI Answer Engine
Abstract
A controlled study of roughly 4,000 trials across three model families — Claude, GPT and Gemini, with Perplexity used to validate the instrument — testing which editorial choices in a machine-facing facts document actually change whether a model cites you. Three findings dominate. Rigor markers on your numbers are the strongest thing you control. The single biggest lever is not on your document at all; it is whether an independent source carries the same data. And an undisclosed commercial conflict is the most damaging signal we measured, collapsing top-two inclusion from 99% to 4%, with a disclosed editorial firewall recovering most of that in the language tests. The study is self-interested and says so: sections 4 and 5 test our own naming and the exact structural conflict we carry, and we publish the numbers that went against us.
A disclosure, up front. This is self-interested research. Sections 4 and 5 test our own naming and brand architecture, and section 5 measures the exact structural problem we have: a research publisher with a commercial practice attached. We publish the numbers that went against us, including the ones that argue we should give things up.
What we set out to measure
We wanted to know what a company should write in the machine-facing facts document it publishes for AI answer engines. We ended up testing two things that mattered as much or more: who publishes it, and what commercial interest the publisher discloses. Every arm used a competitive design in which documents differ only in the lever under test, so an effect is attributable to that lever and not to length, tone, or luck.
The uncomfortable production fact came first. Across 18 days and roughly 15,000 requests from AI-engine crawlers and answer-time agents to a live production site, not one requested /llms.txt. One site, one window. What the answer-time agents did fetch was the homepage, one statistics-dense blog post, and individual product pages. Serve the file, but put load-bearing facts in rendered HTML. The lesson we took: serve the file, but put every load-bearing fact in rendered HTML where the answer-time agents actually read.
What to write: rigor beats persuasion
- Rigor markers are the strongest thing you control. Stating the denominator, the exclusions, the instrument, the metric definition, and a checkable public source beat an otherwise-identical promotional document in essentially every comparison we ran.
- The biggest lever is not on your document at all. The only interventions that cleared the “self-published, unverified” objection were separate sources carrying independent data. Replicated at n=50 per arm on GPT and Gemini, a corroborating source took the objection to 0 of 50 — and, tellingly, so did a contradicting one. What models reward is that someone else is in the record, not that the other source agrees.
Who publishes it
Publishing a versioned measurement standard moved top-two inclusion from 4% to 96%. Renaming the publisher to assert standards-setting authority moved it by 0.15 points. Claiming the authority is worth almost nothing; publishing the artifact it implies is worth almost everything. Measured on document rankings, not on live search results.
Read the two numbers together. A name-and-deed 2×2 showed the document is worth roughly fifteen times the name, with no interaction between them. Claiming to be a standards body is worth almost nothing; publishing the versioned standard the claim implies is worth almost everything. That is why we ship the AI Citation Measurement Standard and grade ourselves against it in public, rather than leaning on the name.
The conflict of interest, measured on ourselves
- Selling compliance with your own standard is the sharpest offence. Selling compliance with a standard you publish collapses top-two inclusion from 99% to 4%. This is the structural problem we have, measured on ourselves. It is why the instrument and the optimization service are separated.
- A disclosed firewall recovers most of it — in the language. An explicit editorial firewall recovers 78% of that collapse (95% CI 70 to 87). Tier 2 evidence. Holds pooled across model families but is NOT statistically significant on GPT alone, and it measures the WORDING of a disclosure, not the existence of an organizational separation. Language moved the ranking; do not read it as proof a firewall works.
- Auditable commitments, not reassuring words. A scarcity pitch, one client per vertical, recovers 8% of the same collapse. What models reward is auditable commitments, not reassuring language.
- The penalty is about substitution, not use. A disclosed commercial tie costs 2.70 ranking points when several publishers report the same finding, and costs nothing when the publisher is the only source for it. The penalty applies to selection among substitutes, not to use. Being the only source for a number protects it; being one of several does not.
A note on fabricated authority
We also measured what happens when a document attributes its figures to an authority that does not exist. We report this as defence, not instruction, and deliberately omit the construction that worked. An invented body paired with a fabricated report number outranked every honest document we tested, and models relayed the fabricated citation to the user without flagging it. Detection ran unreliable in both directions: models also rejected a real trade source as fabricated. The practical takeaway is defensive. A competitor that claims third-party verification which does not resolve — a body with no footprint, a report number that leads nowhere — is a detectable signal, and monitoring for it is part of what we do.
Our policy: the editorial firewall
This paper argues that an unmanaged conflict is the most damaging signal a publisher can carry, so we state our own answer as binding policy rather than sentiment:
AI Citation Institute publishes the AI Citation Measurement Standard and also operates a consulting practice. Research is produced independently of that practice. Clients receive no preferential treatment in measurement and are excluded from our published market rankings; their outcomes appear only with consent or fully anonymized, never ranked against non-clients. The standard is versioned publicly so any change can be audited.
Each clause is a real operating constraint we accept. “Excluded from our published rankings” means a client can never appear in the market Index we rank non-clients on in the scoreboard; their outcomes live in a separate, labeled section, shown only with the client’s consent and otherwise anonymized. Publishing that promise before committing to it would be exactly the failure this paper documents.
Limitations
Every “Gemini” number here is gemini-2.5-flash called via API, not Google’s AI Overviews, which add a retrieval and citation-selection layer we did not test. These effects are measured on document rankings in a competitive design, not on live search results, and the firewall result specifically is Tier-2 evidence: it holds pooled across model families but is not significant on GPT alone, and it measures the wording of a disclosure, not the existence of an organizational separation. One transfer test used a live production site (Gavelist, which shares our ownership — a self-referential test, not a client disclosure). This is a checkpoint in an ongoing program, not a final word. We publish the reading and will re-run it as the panel drifts.
Source: The AI Citation Institute, 'What makes an llms.txt get cited', §5; independently recomputed in the fresh-eyes adversarial review, 2026-07-27 (n=54 rankings, pooled). Independently recomputed in the fresh-eyes adversarial review, 2026-07-27.
Cite this record
Open accessAI Citation Institute. "What Actually Makes a Source Get Cited by an AI Answer Engine." Research record 2026-003, 2026-07-25. https://aicitationinstitute.org/research/what-makes-an-llms-txt-get-cited (CC BY 4.0).
Released under CC BY 4.0 — quote it, chart it, cite it. All we ask is attribution back to this record.
Related: what an llms.txt file is, the voice experiment, the measurement standard, and our methodology.