You do not design leaderboards. You are just an ordinary person surrounded by them: rankings for picking a university, ratings for picking a hospital, scores for picking a restaurant, and even a leaderboard for which AI is strongest. In C02 you saw a seven-lever scorecard, which is a professional tool for designers. Translate the same seven things into seven plain questions, ask them the next time any leaderboard crosses your feed, and you can roughly judge how likely it is to be distorted and how far to discount it.
Behind these seven questions is one causal chain. A number first gets compressed into a comparable rank. That rank then gets published. Once the rank is public, the people ranked start reallocating effort toward it, redefining their work, sometimes outright faking it. These responses pile up, and in the end the world really is reshaped into what the ranking measures. Whether that chain turns, and how violently, is decided in seven places. The first three ask what the ruler looks like. Question one: is there only one leaderboard in this field, or several that say different things? A single authoritative leaderboard is the most dangerous, because there is nowhere to go and nothing to check it against. American law schools recognize essentially only U.S. News, and the effect is extremely strong; business schools have five or six rankings coexisting with contradictory conclusions, so schools can pick whichever one flatters them and tell that story, which dilutes the power of any single leaderboard. Question two: is the rank actually tied to money, to admissions, to someone's livelihood? The tighter the tie, the stronger the motive to fake. Credit ratings are written into bank capital requirements, so pulling one thread moves everything; a most-liveable-cities list looks lively but locks down nobody's money, and the pressure is far lower. Question three: does it give you a single placing, or a set of components? An overall ranking compresses many dimensions into one line, which looks clean but hides how the weights were chosen.
The last four questions ask how the ruler gets used and who is watching. Question four: how often does it update? The more frequent and the closer to real time, the more it works like an unbroken surveillance camera, with those ranked never able to relax, and both the pressure and the maneuvering rise. A live ladder refreshing monthly and an assessment done once a decade force completely different behavior. Question five: is it a rank where you go up only if I go down, or a score everyone can meet at once? A zero-sum rank manufactures an arms race: not pushing forward is the same as falling back, so everyone escalates. Officials competing for promotion on GDP rankings is zero-sum and drives vicious competition; restaurants meeting an absolute hygiene standard is not, and your passing does not stop anyone else passing, which is far gentler. Question six: who reports these numbers, and can they fill them in themselves or reclassify them? If the leaderboard runs on data self-reported by those being ranked, and lets them reclassify the unflattering parts out of sight, there is a lot of room to go around it. U.S. News long relied on schools reporting their own figures, and schools put low-scoring students into programs that were not counted; whereas code that actually runs and a result that reproduces cannot be inflated by self-report. Question seven: do the people being ranked know the scoring rules, and will they optimize for them? Once they have the formula figured out and tune specifically for it, the leaderboard distorts fast. Model vendors studying benchmarks and targeting their scores, and law school deans reverse-engineering every weight in U.S. News, are the real-world version of this question.
Do not swing to the other extreme: not every leaderboard deserves a discount. A leaderboard that scores low on these seven questions has mild reactivity, and can even be beneficial. The economists Jin and Leslie studied restaurant hygiene grade cards in 2003, and they score low on nearly every question: not zero-sum, not real time, consequences not severe, judged on an absolute standard rather than a rank. The result was that restaurants really did clean up their kitchens, foodborne illness fell, and no arms race followed. So the seven questions are not for beating every ranking to death. They help you tell whether the one in front of you is benign or high risk.
Score them in your head one by one, and the more high scores, the more this leaderboard deserves a discount. Three closing lines to keep. First, accurate is not good: a ranking can be very accurate and still worthless, because what it measures is the world it made. Second, improving the metric is not removing reactivity, so do not believe in a perfect one-and-done leaderboard. Third, the most dangerous metric is usually not an obviously terrible one, but a "good enough" one. It looks reliable, so it wins trust first. Then it gets bound to high stakes. Once everyone piles on to optimize it, it quietly fails in the process.
One open question: these seven questions help you identify how easily a leaderboard can be distorted, but they still cannot tell you how much credibility is left after the distortion, meaning how large a discount is right. Turning qualitative identification into a quantitative discount is a problem the field has not solved.
The one-line takeaway: for any leaderboard, ask these seven things first: how many leaderboards, tied to money or not, how blunt, how often, zero-sum or not, self-reported or not, and whether those ranked will tune themselves to it.
Sources / further reading
- Seven-lever framework synthesized from
00-SYNTHESIS-总纲.md §4(causal chain plus seven scorable levers) and the C02 section ofresearch/14(operationalized questions, scoring examples, reader checklist). - Monopoly as a moderating variable: Sauder & Espeland (2006). "Strength in Numbers?" Indiana Law Journal 81(1) (single-leaderboard law schools vs multi-leaderboard business schools).
- Arbitrariness of weights in composite indices: Saisana, M. et al. (2005). "Uncertainty and sensitivity analysis techniques as tools for the quality assessment of composite indicators." JRSS A 168(2).
- The three corollaries ("accurate ≠ good," "improving the metric ≠ removing reactivity," "the most dangerous metric is the good-enough one") are in
00-SYNTHESIS-总纲.md §4; new evidence on continued forced ranking after withdrawal and on commercialized judges is inresearch/deep/D1§1.4.