EST. MMXXVI · THE INSTITUTION OF RECORD FOR AI CITATION · OPEN METHODOLOGY

The AI Citation Institute sealTHE AI CITATION INSTITUTEEx corpore, auctoritas.

The AI Citation Institute · The Answers

How do AI engines choose citations?

AI engines choose citations in two stages. First is retrieval: the engine pulls a set of candidate pages from its index or a live search, favoring pages that are indexed at all, whose title matches the query, and that are fresh. Then comes selection: from those candidates, the model quotes the one page that hands it a passage it can lift whole. Selection rewards three things — a standalone, extractable block that answers the exact question without needing surrounding context, specific first-party numbers a model can't invent on its own, and outbound citations to sources, which lend the page borrowed credibility. The strongest on-page lever we've measured is neither of those, though: inbound internal links. In our 91-post, four-engine audit, a page with zero inbound links was self-cited about 4.5% of the time; a page with ten or more was cited about 44% — same content, different link graph, roughly a tenfold swing. So getting quoted is less about writing for the model and more about being retrievable first, then giving the model the cleanest thing on the page to grab.

Two stages: retrieval, then selection

It helps to split the pipeline. Retrieval decides which pages are even in the running — an engine can't quote a page it never pulled. Selection decides which of the retrieved pages actually gets the citation. Most "my page is great but never cited" cases are really retrieval failures: the page was never a candidate. The two stages reward different things, so it's worth knowing which one is failing.

Signals that help retrieval versus signals that win selection
StageWhat it rewardsWhy
retrievalIndexation, title–query match, freshnessThe engine has to find the page and consider it relevant before it can be quoted. A crisp title that mirrors the query and a recently updated page both raise the odds of being pulled.
selectionA standalone extractable passage that answers the question directlyThe model prefers a block it can lift whole, without stitching in surrounding context. Answer-first paragraphs beat buried answers.
selectionSpecific first-party numbersConcrete figures a model can't invent give it something to attribute. Vague claims get paraphrased; exact numbers get quoted and cited back to you.
selectionCited sourcesOutbound citations lend a page borrowed credibility, so the model treats it as a place to source from rather than just another opinion.

The strongest lever we've measured: inbound internal links

If you only change one thing, change the link graph. In our 91-post, four-engine audit (2026-07-11), a page with zero inbound internal links was self-cited about 4.5% of the time. A page with ten or more inbound links from other pages on the same site was cited about 44% — roughly a tenfold swing on the same content. Nothing else we tested moves the number that far. The mechanism is retrieval: inbound links are how a page gets crawled, indexed, and treated as central rather than orphaned. Write the extractable passage, but make sure the page is well-linked so it becomes a candidate in the first place.

Figures from our 91-post, four-engine audit, 2026-07-11. Same content, different link graph, roughly a tenfold swing.

Freshness helps retrieval — but only real changes count

Freshness is a genuine retrieval signal, not a myth. Under our v5 rubric, content updated within 30 days with a substantive change earned about 3.2x more ChatGPT citations than stale content. The catch is that date-only bumps don't count — editing the page's timestamp without changing what it says buys nothing, and bulk re-bumping can look like manipulation. Update pages when you actually have something new to say.

Engines weight these signals differently

The same page performs differently across engines because each one grounds and gates differently. Gemini leans on Google's grounding and cited us most. ChatGPT (Bing browse) has the weakest brand-authority gate, which makes it the most reachable engine for a brand below the authority threshold. Perplexity rewards citation density and tight title match. We read these on a frozen five-engine panel, each query sampled k times, so the rates are comparisons of the same instrument rather than one-off screenshots.

Observed per-engine citation behavior
EngineGroundingObserved rate / tendency
GeminiGoogle-grounded~63% — cited us most
ChatGPTBing browse~47% — weakest brand-authority gate, most reachable for below-threshold brands
PerplexityLive search43–71% — rewards citation density and title match

Observed Gavelist rates on our frozen five-engine panel, each query sampled k times. See /methodology for the instrument.

Common questions

What's the single biggest lever to get cited by AI?+

Inbound internal links. In our 91-post, four-engine audit, a page with zero inbound links was self-cited about 4.5% of the time; a page with ten or more was cited about 44% — a roughly tenfold swing on the same content. It works mainly by improving retrieval: linked pages get crawled, indexed, and treated as central.

Does adding FAQPage schema or an llms.txt file get me cited?+

Those are hygiene, not levers. Schema helps a machine parse a page it has already retrieved, and llms.txt is a courtesy file. Neither reliably moves citation rates on its own. Spend the effort on the link graph, an extractable answer-first passage, and real freshness instead.

Why does one engine cite my page and another ignores it?+

Each engine grounds and gates differently. Gemini leans on Google's grounding, Perplexity rewards citation density and title match, and ChatGPT (Bing browse) has the weakest brand-authority gate, so it's the most reachable for a brand below the authority threshold. The same page can clear one engine's bar and miss another's.

keep reading

See what AI says about your brand right now — free, on a live engine.

Scan your domain