Where we fall short
Of 16 clauses we wrote, 6 are not fully met: 0 we do not meet at all and 6 we meet in part. Most of the partials are reporting gaps rather than instrument gaps: the measurement is made correctly and then published less completely than this document requires.
One of them is a boundary rather than a gap, and it will not close. Since 2026-07-30 every reading records the model that actually answered it. Every reading before that date is permanently unattributable, because the log we keep of instrument changes records that a shift happened and not which model was running on either side of it. Those readings store no model, and no model is the correct final answer for them. Supplying one later would be inventing the history the clause is there to protect. An instrument that can say where its own memory ends is worth more than one that cannot.
Two more clauses failed when we first checked, on the day this page was written, and were fixed before it shipped. A panel question could be retired with no record of who did it, and we described our own aggregation as a median when the code takes a majority vote. Neither was found by reading our documentation, which asserted both correctly. They were found by requiring code as evidence.
- P1 we partly meet this The query panel is fixed and published
- P3 we partly meet this Branded and unbranded are reported separately
- R2 we partly meet this Name the model, not just the engine
- E1 we partly meet this Observed and projected are visually distinct
- E2 we partly meet this Causal claims carry a grade
- D2 we partly meet this Publish your own reading, including when it is bad
Using this
Adopt it, adapt it, or cite it to argue with us. If you measure differently and can show why, that is more useful to us than agreement: a standard nobody contests is usually one nobody read. Clause identifiers are stable, so S2 will mean the sampling floor in every future version.
Version 1.0, published 2026-07-30. Conformance verified 2026-07-30. The instrument this was written against is described in full on the methodology page, and the research behind the measured clauses is on the research page.