There is a movement just getting started: evaluating evaluation itself. The diseases of AI evaluation have been laid out one by one: contamination, saturation, leaderboard gaming, sandbagging. Following the chart down, the natural questions are who is doing the repairs, with what, and whether the repairers themselves can be trusted. The answer has three layers: a group of researchers is trying to upgrade benchmark-building from a craft into a discipline with standards; meanwhile the most influential judge is turning into a business; and the variable of "does the judge have interests of its own" is something the field has only recently begun to look at squarely.
The movement starts from a measurement-theory concept: construct validity. Plainly: does your ruler measure the thing you meant to measure? A bathroom scale measuring "health" and a vocabulary test measuring "intelligence" both substitute a measurable proxy for an unmeasurable target, and the crack in between is the validity problem. Raji et al. formally imported this concept into AI evaluation in 2021, pointing out that benchmarks like ImageNet and GLUE, treated as rulers of "general capability", are products of specific data, specific metrics and specific annotation habits. A product like that can only ever hold the particular choices baked into it at the start, and structurally it cannot carry the word "general". Using a narrow ruler as a universal one is itself a measurement accident. From then on, "what does this benchmark actually measure" became a question that could be asked seriously for the first time.
What turned the critique into a physical exam report is BetterBench. Stanford's Reuel et al. in 2024 listed 46 best practices for "what a competent benchmark should do", covering the whole life cycle from design and implementation to documentation and retirement, then scored 24 mainstream benchmarks item by item. The result: quality varies enormously, and most benchmarks neither report statistical significance (without a test you cannot tell whether a score gap is a real difference or sampling luck) nor guarantee that anyone else can reproduce their results. Statistical significance testing and reproducibility are exactly the two most basic thresholds of experimental science. The diagnosis compresses to one sentence: evals are experiments, yet they lack the basic standards of experimental science. The whole industry is measuring the most expensive things with rulers that never passed factory inspection.
The accompanying documentation proposal follows machine learning's existing traditions: models have a model card (a manual: what it can do, what it cannot, where it falls over), datasets have a datasheet, and so benchmarks should have an eval card: stating which construct it means to measure, the scope where it applies and where it does not, its statistical conventions, its known defects and its life cycle status (still fresh, or already ground to saturation). Let "what conclusions this score can support" travel with the score as a manual.
Institutionalisation came thick and fast between 2024 and 2026. Academia published the programmatic article Toward an Evaluation Science for Generative AI Systems; Anthropic published a blog post enumerating the practical difficulties of evaluation while Evan Miller wrote Adding Error Bars to Evals, importing first-year statistics tools like error bars (the small line next to a score showing the range of uncertainty) and clustered standard errors into evals one by one, and the paper contains a line that could serve as the movement's epigraph: an empirical science is only as good as its measurement tools. Anthropic also funded third-party evaluation directly. On the academic side, Hardt's 2026 The Emerging Science of Machine Learning Benchmarks simply makes the benchmark itself the object of study. There is an isomorphism among these prescriptions: putting error bars on scores and the antidote to B13's optimiser's curse (disciplined scepticism about a top score you selected, shrinking it back) are two ways of writing the same medicine.
The movement's other face is the commercialisation of the judge, the thread Y11 planted at its end. LMArena incorporated out of its Berkeley academic project: a $100 million seed round in May 2025 (valuation about $600 million, led by a16z), a further $150 million Series A in January 2026 at a valuation jumping to $1.7 billion, alongside the launch of segmented leaderboards and paid evaluation services, making the ranking formally a product. The governance question takes shape with it. Style control proves the operator is capable of self-correction (explicitly subtracting the style that got optimised into existence, as Y11 covered); but when the operator is a company monetising the leaderboard's credibility, whether the will to self-correct persists, and whether it tilts towards large customers, is something no external audit can establish. This is not an accusation of impure motives, it is a description of a structure: when the judge's valuation and the judged parties' performance sit on the same balance sheet, "the judge is neutral" is demoted from a default assumption to a claim awaiting verification. The FrontierMath funding controversy in Y14 is another cross-section of the same structure.
Following that structure, a new lever can be proposed for C02. C02 gathers the whole course into a scorecard: seven levers (monopoly, stakes coupling, bluntness, release cadence, zero-sumness, available bypass channels, the reflexive capability of the measured), each scored, predicting how badly a measurement system will distort. The evidence of 2025 to 2026 points at an unlisted candidate: the judge's interest exposure. It asks whether the scoring body's revenue, valuation or data supply depends on the parties being scored. LMArena becoming a unicorn, FrontierMath's funder holding exclusive access, and the curated benchmark selections at vendor launch events (the same model's score can drift several percentage points on peripheral settings, and which benchmarks go on the model card is itself a performance) all fall on this lever. It is orthogonal to the first seven: a monopolistic, high-stakes, continuously updated leaderboard can have a neutral judge or a non-neutral one, and the first seven levers cannot measure the difference. The operational question can be very plain: what does the judge lose by reaching a conclusion unfavourable to its funder? The larger the loss, the more the reading should be discounted.
Why is this ground empty right now? The returns to building models are direct, enormous and yours; the returns to building evaluations are indirect, slow and mostly other people's, since a good benchmark is a public good that everyone can use and nobody has to pay to maintain. So the people gaming a measure will always outnumber the people repairing it. The industry has said it out loud: the hardest thing is not building models, it is building evaluations. Multiple organisations turning towards real tasks and expert-designed evaluations, and Anthropic funding third-party evaluation, amount to subsidising the production of a public good. For you, reading this far, the empty ground is the opportunity: what this direction lacks right now is not another model that scores well, it is people who take measurement seriously.
One open question: can evaluation science avoid catching the disease it diagnoses? Once the 46 best practices become an industry standard, benchmarks optimised against the checklist will appear; once the eval card is tied to reputation, it will be written as marketing copy. Whoever writes the rules for judges is also a judge who can be Goodharted. This recursion is handled head-on in G06 (who audits the auditors).
The one-line takeaway: evaluation itself needs evaluating; and the fastest question about whether a score deserves belief is what the scorer does for a living.
Sources / further reading
- Raji et al. (2021). "AI and the Everything in the Whole Wide World Benchmark." arXiv:2111.15366 (NeurIPS D&B, the construct validity critique).
- Reuel et al. (2024). "BetterBench." arXiv:2411.12990 (NeurIPS Spotlight, 46 best practices, 24 benchmarks audited).
- "Toward an Evaluation Science for Generative AI Systems." arXiv:2503.05336; Miller, E. (2024). "Adding Error Bars to Evals." arXiv:2411.00640; Anthropic. "Challenges in evaluating AI systems" (blog).
- Hardt, M. (2026). "The Emerging Science of Machine Learning Benchmarks." (mlbenchmarks.org).
- LMArena commercialisation: Reuters 2026-01-06 (Series A, $1.7bn valuation); Singh et al. (2025). "The Leaderboard Illusion." arXiv:2504.20879.
- "Humanity's Last Exam" (2025). arXiv:2501.14249; Glazer et al. (2024). "FrontierMath." arXiv:2411.04872; Wang et al. (2024). "MMLU-Pro." arXiv:2406.01574.
- Deeper treatment in
research/05§1, §7;research/deep/D4sub-lines 3 and 4;research/deep/D5.