The seven levers are a scorecard. Take any ranking or evaluation, score it item by item, and you get an estimate of how badly it will be distorted. You learned reactivity in R01, Goodhart's law in B01, and reward hacking in Y04, and all three describe the same thing, but all three stop at the qualitative judgment that "it will distort." What this scorecard does is turn that judgment into a quantity you can compare across cases: which system is more dangerous, and which way a given system is drifting over time.
It is not an empirically calibrated absolute index, so do not treat the final number as having measured anything. It is a structured comb, forcing you to break the vague worry of "will this leaderboard get broken" into seven concrete questions you can answer separately.
Here is what each of the seven levers asks:
- Monopoly: how many independent rankings cover this object? One authoritative leaderboard gives the strongest disciplinary force; four or five coexisting and contradicting each other gives the weakest.
- Stakes: is the number tied to money, to life and death, or to a release decision? The tighter the binding, the fiercer the response.
- Bluntness: is the output a total ordering, or a multi-dimensional dashboard? Compressed into a single rank, the illusion of discrimination is strongest.
- Cadence: how often is it published? A continuous live ladder manufactures permanent readiness for battle.
- Zero-sum: is it a relative rank or an absolute standard? When your rise requires my fall, an arms race follows.
- Bypass routes: how long is the causal chain between metric and real goal, and how many side roads run around it? Self-reporting and reclassification make it easy to get around.
- Subject reflexivity: will the party being measured reverse-engineer your formula, and is it itself an optimizer?
Flip through a few presets and the pattern emerges. The U.S. News law school ranking pushes nearly every lever to the maximum: sole authority, hard binding to admissions and budgets, compressed into a total ordering, published annually, zero-sum, bypassable through self-reported data, deans working full time to reverse-engineer the formula. It takes all twenty-one points across the seven items, and it reaches that ceiling, so it is the extreme specimen of reactivity. This maximum-score specimen is exactly the sample in Espeland and Sauder's fieldwork that could "barely be buffered at all." Business school rankings give a different number. Because five rankings coexist, the monopoly cell drops from 3 to 1. With monopoly down, the overall score falls a tier along with it. This lower score lines up exactly with the finding in Sauder and Espeland's Strength in Numbers: ranking effects on business schools are noticeably weaker. Plural measures dissolving each other is the most usable half of this regularity.
Add a reading from the AI side. A live ladder like Chatbot Arena has medium monopoly, high stakes, high bluntness, real-time cadence, zero-sum structure, and subjects that are themselves optimizers. The seven together land high. Once the score is high, its Goodharting flares up on a scale of months, an order of magnitude faster than university rankings. Speed is not because AI is fonder of cheating. It is because the subject is itself an optimizer, a structural difference rather than a moral one. Vote-manipulation research supplies a quantitative footnote: simulated on real historical votes, about 20,000 votes can push a model roughly ten places up the ranking; and getting hold of even a small amount of Arena data and overfitting to its particular distribution can deliver up to a 112% relative performance gain, which need not be a gain in general quality at all.
Now a benign contrast. Los Angeles restaurant hygiene grade cards use only three tiers, A, B and C, with no total ordering, soft consequences, and subjects with almost no ability to reverse-engineer anything, so the seven items add up to single digits. What Jin and Leslie measured in 2003 was real improvement in hygiene and a drop in foodborne-illness hospitalizations, with no arms race. The same reactivity mechanism, with the levers low, can push quality upward. Which shows what the scorecard measures is not whether reactivity exists, but which way reactivity will push.
The scorecard also brings out three less obvious conclusions.
First, reactivity can simultaneously raise measurement validity and destroy measurement value. The self-fulfilling channel makes a ranking "more accurate," because the world has already been remade to match it. But what it measures is already the world it built itself; "accurate" and "good" therefore become two different things.
Second, improving the metric is not the same as removing reactivity. Every revision of a metric is itself an object of gaming: U.S. News tweaks its formula every year, Arena added style control, and vendors turned around to optimize the residual left after style control. A metric and its subjects are a co-evolving pair. You move, it follows.
Third, the most dangerous metric is not a bad one. It is a "good enough" one. A highly correlated proxy earns trust, trust gets it bound to stakes, stakes bring strong optimization, and under strong optimization it fails. Gao's overoptimization curve, rising then falling, is exactly the fate of a "good enough" metric.
Evidence from 2025 and 2026 has forced an eighth candidate lever out of this seven-cell scorecard. You have seen its name in Y15: judge's stake. It asks something none of the first seven can reach, namely whether the scoring body's revenue, valuation or data supply depends on the party being scored.
Three contemporary events land on this lever. First, the most influential live leaderboard, LMArena, went from a Berkeley academic project to a company, raising a $100 million seed round in May 2025 (at a valuation of about $600 million), then another $150 million Series A in January 2026 at a valuation jumping to $1.7 billion, while turning the ranking into a paid product. When the judge is itself a company monetizing the credibility of its leaderboard, whether its willingness to correct bias persists, and whether it tilts toward large customers, is not something any external audit can establish. Style control proves the operator is able to explicitly subtract out optimized-for style. Ability is not willingness. Second, the contamination-resistant benchmark FrontierMath was funded by OpenAI, and OpenAI also obtained exclusive access to the great majority of the hardest problems, then used that set to announce o3 scoring 25% (the previous generation managed only 2%). The selling point that "the model has not seen the questions" thereby degrades from a verifiable fact into a verbal assurance you can only choose to believe or not. Third, benchmark cherry-picking at vendor launches: the same model's scores can drift several percentage points with peripheral settings, and which benchmarks go on the model card is itself a performance.
This lever is orthogonal to the first seven. A monopolistic, high-stakes, real-time leaderboard can have a neutral judge or a compromised one, and the first seven cannot tell the difference. The operational form of the question can be very plain: if the judge published a conclusion unfavorable to whoever pays it, how much would it lose? The larger the loss, the more you should discount the reading.
One open question: whether it is seven levers or eight, adding them with equal weight is itself an uncalibrated choice. Could a large body of real cases be used to actually estimate each lever's weight and the interactions among them, upgrading this scorecard from "teaching intuition" to "a verifiable index of reactivity strength"? That is a patch of open ground nobody in the field has systematically claimed.
The one-line takeaway: to predict how badly a leaderboard will be broken, do not guess, score it item by item; and remember to ask whether the scorer has money on the outcome.
Sources / further reading
- Framework synthesized from
00-SYNTHESIS-总纲.md §4. Arbitrariness of composite indicators: Saisana, Saltelli & Tarantola (2005). JRSS-A 168(2):307–323. - Monopoly and dissolution by plurality: Sauder & Espeland (2006). "Strength in Numbers"; benign reactivity: Jin & Leslie (2003). QJE 118(2) (restaurant hygiene grade cards).
- Eighth lever (judge's stake) and new evidence: LMArena funding, Reuters 2026-01-06; FrontierMath funding controversy, TechCrunch 2025-01; see
Y15andresearch/deep/D4sub-line four,research/deep/D5.