EST. MMXXVI · THE INSTITUTION OF RECORD FOR AI CITATION · OPEN METHODOLOGY

The AI Citation Institute sealTHE AI CITATION INSTITUTEEx corpore, auctoritas.
← all findings
11 min read

We published all 37 measurements we ever took of our own visibility — including the ones that went down, the parser bug, and the re-scores.

For fifteen weeks we have measured one business — our own auction-software product — on the same AI-visibility instrument we sell. Here is the entire run ledger: every completed scan, the two that failed, the week a parser bug inflated a number, the platform shifts that moved every brand at once, and how we handle a panel that keeps growing. Nothing retouched.

Most case studies show you a line that goes up and to the right. They pick the metric that flatters, the window that flatters, and the endpoint that flatters, and they leave the failures on the cutting-room floor. This post is the opposite exercise. We have been measuring one business on a fixed AI-visibility instrument since 2026-04-23 — our own auction-cataloging product, Gavelist, which we call Client Zero because it is the property we own and can publish without anyone's permission. Below is the complete ledger of every scan we ever ran on it: the climbs, the drops, the two runs that failed outright, the week a parsing bug inflated a number, and the platform-wide shifts that moved every brand we track on the same day. We publish the whole thing because a measurement company that only shows you its good weeks is not a measurement company.

One discipline governs everything that follows, and it is the single most important thing on this page. We do not read a trend across the series and attribute it to our work. The instrument changed under this series more than once — the question panel grew, engines were added, two probes were re-instrumented — and each of those changes moves the number for reasons that have nothing to do with visibility. So the only causal language we allow ourselves is a contrast measured inside a single week, where the instrument was identical on both sides. Everywhere else, we show you the shape and we name the reason the ruler moved. If that sounds pedantic, it is exactly the pedantry that separates a measurement from a marketing chart.

What the instrument is

The instrument runs a fixed panel of buyer questions against five engines — ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews — and reads which brands each answer names. The headline metric is the mention rate: the share of panel answers that name the business. We separate branded queries (which ask about the business by name) from discovery queries (category questions where the business should be a candidate but is not named in the prompt), because those two numbers behave completely differently and conflating them is how people fool themselves. The composite score folds mention, active recommendation, and a head-to-head margin against the nearest tracked rival into a single 0–100 number. The full definitions live on our standard and methodology pages, and the running numbers — including this one — are public on the scoreboard.

The record, week by week

Here is the clean series: 37 completed runs, aggregated to weekly medians, from the first panel run through the most recent completed scan. Two more completed runs exist in our database and are deliberately excluded — they persisted a zero score with no rival named, the signature of a pipeline failure recorded as a success; including them would fabricate a June crash that never happened. We note them here rather than silently drop them. The mention column is the share of answers naming Gavelist; the rival column is its nearest tracked competitor, AuctionWriter; the panel column is how many questions were on the exam that week.

Week ofPanelCompositeMention %Rival (AuctionWriter) %What moved the ruler
2026-04-202016.713.324.4Baseline. Behind the rival.
2026-04-272030.431.333.8Probe-mode change on 04-28 (instrument, not a win)
2026-05-042031.433.836.9
2026-05-112032.833.840.0
2026-05-182033.834.437.6
2026-05-252041.547.541.3First week ahead of the rival
2026-06-012046.950.735.1
2026-06-082046.450.136.3
2026-06-152048.260.052.0Claude probe re-instrumented (10→100 draws)
2026-06-222052.765.352.0Peak weekly mention
2026-06-292552.264.854.4Panel grew 20→25
2026-07-063847.655.340.5Panel grew 25→38 (harder exam; part of the dip is dilution)
2026-07-133848.157.838.0Parser fix landed 07-19 (see below)
2026-07-203842.053.042.6Two all-engine platform shifts (07-19, 07-27)
2026-08-034846.759.541.4Panel grew to 48; latest completed scan (run 2026-08-09, k=1 benchmark — see below)

Read that table honestly and you will see it is not a clean climb. Mention rose from 13.3% to a 65.3% peak in late June, then sat in the mid-to-high 50s through July and reads 59.5% at the most recent completed scan on 2026-08-09. The composite tells a flatter story than mention does, and the gap between them is the point: the composite folds in an active-recommendation term that fell across the platform shifts even as raw mentions held. We could have shown you only the mention line and only through June 22. We are showing you all of it.

The single most important caveat: the exam kept getting harder

Look at the panel column. It starts at 20 questions and ends at 48. We added questions as we discovered new ways buyers actually ask about the category — objection-frontier questions about manual-review trust, house-style templates, batch-and-sync workflows, and more. Every question we add is a fresh chance to not be mentioned, so a growing panel mechanically pushes the mention rate down unless visibility is improving fast enough to offset it. This means a good chunk of every downswing in the table above is the exam widening, not visibility declining. It also means you cannot compare the composite score across a panel change and call the difference performance. When we want a comparison the panel growth does not contaminate, we use the frozen panel below.

The frozen panel: the only apples-to-apples series we have

To get a series where nothing moves except the answers, we freeze everything else. The frozen panel uses only the original 20 questions from the 2026-04-23 baseline (all 20 still run every week; later additions are simply excluded here), and only the two engines — Gemini and Perplexity — that had zero instrument events across the entire record. Mention is presence-based, so the parser fix cannot touch it. This is the series we treat as the headline, because it is the only one where a week-over-week move is unambiguously the answers changing rather than the ruler changing.

Week ofMention % (frozen)Recommendation % (frozen)Composite (frozen)
2026-04-2016.43.814.9
2026-04-2727.811.924.0
2026-05-0435.822.531.7
2026-05-1134.218.328.8
2026-05-1840.023.835.6
2026-05-2562.522.548.6
2026-06-0167.525.054.8
2026-06-0865.030.054.5
2026-06-1560.017.547.3
2026-06-2262.510.047.3
2026-06-2962.527.552.1
2026-07-0680.037.568.8
2026-07-1376.323.861.0
2026-07-2075.020.056.5
2026-07-2780.035.064.0

On identical questions asked of identical engines, mention rose from 16.4% to 80.0% and has held at or above 60% since the week of 2026-05-25. That is the cleanest read of visibility improving that the record contains, and it is the one number on this page we would defend as close to instrument-free. Recommendation is a different and more sobering story, which we get to next. Four rows of this table were corrected on 2026-08-01 when we found that hand-written SQL had let two partial runs (a coverage the frozen-panel code is built to reject) into a couple of early weeks; regenerating from the code moved two composites by 0.1 and left the narrative unchanged. We mention the correction because that is the policy — a corrected number is routine, not an embarrassment.

The week a bug inflated our own number

On 2026-07-19, three minutes before that day's run, we shipped a fix to our answer parser. Before the fix, an entire bulleted list of vendors was scored as a single unit, so a praise word appearing anywhere in the list credited every brand in it — which inflated our stored recommendation rate. Concretely, on the 07-12 core panel our stored recommendation read 53.3%; re-parsed under the corrected logic it was 43.3%, a full ten points lower. That is our own number, moving down, because our own tool had been too generous to us. Every recommendation figure on this page is the re-parsed one. We keep the inflated stored values in the database for audit continuity, but we never chart them, and we would rather tell you the bug existed than quietly serve you the corrected series as if it had always been right.

The parser bug made us look better than we were. We found it, re-scored every affected answer downward, and published the smaller number. That is the whole business model in one sentence.

The two weeks every brand moved at once — a within-week contrast

Twice in nine days — on 2026-07-19 and again on 2026-07-27 — a canary set moved on nearly every engine on the same day, and it moved for every brand we track, not just us. When endorsements withdraw across the entire field simultaneously, that is the platforms changing how they answer, not any one site winning or losing. So here is where the within-week discipline earns its keep. In the week of 2026-07-19, our frozen-panel recommendation fell to 20.0% — but so did the recommendation counts of every tracked brand, together. In that same window, the number of times engines cited Gavelist's own pages actually doubled, from 56 to 131. Read those two facts side by side, inside the one week where the instrument was fixed: retrieval strengthened while endorsement conversion fell across the whole market. We are allowed to say that, because both halves were measured on the same instrument in the same week. We are not allowed to say 'our recommendation rate fell because of X' as a trend, and we do not.

The recovery week is the mirror image, and it is where we are most careful. In the week of 2026-07-27, recommendation came back to 35.0% on the frozen panel. Was that us or the platforms? We ran the same all-brands-together test that justified calling 07-19 a platform event, and this time it gave a different answer. On the shared-query basis, our recommendation share rose about 9.9 points while the rest of the field's gains clustered between roughly 0 and 4 points and one tracked brand actually fell. That is not the uniform, everyone-moves-together signature of a pure platform event. So the honest ruling is: the recovery is part platform and part site, and we refuse to present it as fully earned or as fully external. That is a less satisfying sentence than 'we won,' and it is the true one.

The runs that failed, and the run that lives only in the record

Two core runs in our database carry a failed status, and we are not hiding them: they hit a per-run cost ceiling partway through and never completed a full panel, so they are excluded from every series above. Separately — and this is a subtler disclosure — the completed run we conducted on 2026-08-01, which read 63.2% mention on a 43-question panel, does not exist as a completed row in our current production database; it was regenerated into our dated visibility record during a data migration and its figures stand on that dated record, not on a live database row. We flag that provenance every time we cite it, including here, because 'the number is real but its database row did not survive a migration' is exactly the kind of thing a measurement company should say out loud rather than paper over. Where a figure is post-migration, we verify it against the live database; the 2026-08-09 reading of 59.5% on 48 queries in this post was re-verified against the production database on 2026-08-11.

A note on how these last two numbers were measured

One more disclosure, because it is the kind that quietly breaks comparisons if you skip it. High-intent queries on our panel are normally sampled three times per engine and scored by majority vote, to smooth out the run-to-run variation that AI answers exhibit. The 2026-08-09 benchmark run was executed with that repetition turned down to a single sample per query, by deliberate override, to control cost while the scan cadence was paused. That does not make the 59.5% wrong, but it does mean it was measured under a slightly noisier setting than the scheduled runs it sits next to in the table, and a difference of a few points across that boundary could be sampling noise rather than a real move. We would rather you know that than infer a trend we cannot support.

How this could be wrong

This is one business, our own, measured through engine API surfaces rather than the consumer apps — API-surface rates can differ from what a person sees in the app. It is a single case, not a typical result and not a guarantee. The composite score folds several sub-metrics whose weights we chose; we publish the formula so you can disagree with it. Small panels move in coarse steps, which is why we report counts wherever we have them. And the whole series sits on top of engines that change weekly, so any single week's reading is a fact about that week, not a stable property of the world. The frozen panel controls for our own instrument changes; it does not control for the engines drifting underneath everyone.

Why publish the drops at all

Because the value of a measurement is entirely in whether you can trust the measurer, and the only way to earn that is to show the measurements that did not flatter us. We publish the parser bug because it inflated our own number. We publish the failed runs because they are part of the honest count. We publish the platform shifts because they explain moves we would otherwise be tempted to take credit for. We publish the panel growth because it makes our own numbers look worse. This is the same standard we hold every page in this research program to, and the same one we point at the businesses we measure: dated observations, a stated method, disclosed limitations, and a correction policy that treats an updated reading as routine. The full case study for this engagement, with the regime-annotated chart, is published; the definitions and correction policy are on our standard; and the live numbers, ours included, are on the scoreboard. If a future scan changes what is true here, this page will change with it.

Sources

  1. Gavelist — live case study (Client Zero)
  2. The measurement standard (definitions, correction policy)
  3. How we measure (the instrument, live panel)

Every external figure above was verified against the primary source before publication. Our own figures come from the instrument described on the methodology page; the claims ledger for this post is part of our research record.

Cite this article

Open access

AI Citation Institute, "We published all 37 measurements we ever took of our own visibility — including the ones that went down, the parser bug, and the re-scores.", 2026-08-11. https://aicitationinstitute.org/blog/we-published-every-scan-we-ever-ran-on-ourselves (CC BY 4.0).

Quote it, chart it, cite it. All we ask is attribution back to this page.

This is exactly the kind of first-party finding we build client content around — and measure on a frozen five-engine panel.