Technical measures fix the ruler itself. One level above sits the rules governing whoever holds the ruler. When hidden items, pairings and randomization get routed around one by one, the last line of defence is institutional: who audits the auditors, whether the evaluator can be funded by the evaluated, whether the ranking has transparency and an appeals route. There are four institutional defences. They are the last layer, and also the most expensive and slowest one.
You saw in R13 that two years after top law schools collectively withdrew from US News, the ranking carried on regardless using public data, untouched. Reactivity cannot be stopped by the tested party simply refusing to cooperate. G02 covered decoupling measurement from judgement. When the ruler gets exploited no matter how you fix it, can institutions catch what falls through?
The first piece of institutional defence, and the most critical one, is audit independence: whoever evaluates you must not live on your money. Counterexamples are everywhere. FrontierMath, a contamination-resistant math benchmark that should have been the most trustworthy, turned out to be funded by an AI company. That funder had access to the great majority of the hardest problems and their answers. Holding the answers, it turned around and used the same set of problems to announce its own model's score, and the six mathematicians who wrote the items were not told in advance. University rankings are the same. QS ranks universities while selling consulting and paid ratings to those same universities. Research found that 22 of the 28 Russian universities in the main table bought roughly 2.85 million dollars of its services over eight years, and the frequent buyers rose on average about 140 places more than the ones that did not buy. Even US News has been called out: boycotted by universities on one side, while making money selling those same universities paid badges saying where they rank in US News. All of this shares exactly one thing: a financial tie between evaluator and evaluated, and independence is gone.
The harder the consequences, the harsher the capture, and the World Bank's Doing Business ranking is the most glaring example. That table directly shaped national reform agendas, and India's prime minister publicly demanded the country rise 100 places. Precisely because the consequences were large, the ranked parties had a strong incentive to capture the scoring itself. In 2017, Chinese officials applied repeated pressure. The World Bank's then-CEO took personal charge and instructed the team to find "adjustable" data points, and changes eventually made in several ambiguous areas lifted China from 85th back to 78th. That move happened to land right in the middle of a 13 billion dollar capital increase negotiation, in which China's shareholding was about to rise from 4.68% to 6.01%. After the scandal broke, the Bank discontinued the report permanently in 2021 and replaced it with a new one called B-READY, claimed to have a more transparent methodology. Same medicine in a new bottle: at the end of 2025 Hong Kong was placed in the top 20, and the SAR government immediately said publicly that the rating was outdated and unfair. This confirms the judgement in R13 and exposes a through-line: switching to a cleaner method does not dissolve reactivity, and as long as the consequences are large, the ranked parties will still contest, challenge and lobby. What is really missing is governance, not just methodology.
Governing a ranking takes three companion pieces. First, methodological transparency: publish the weights, publish the algorithm, and publish how much the rankings would shift under an equally reasonable alternative algorithm. Do not treat the ranking as an unquestionable black box. Second, an appeals mechanism: the evaluated party needs a channel to challenge and correct, rather than only swallowing a wrong number. Third, a plural ecosystem of tables: do not let one player dominate. Business schools have five or six coexisting tables with different methodologies, so a school can pick the one that flatters it for its story, and the dominance of any single table is diluted by the others. That is exactly the logic in R01 where multiple rankings restore room for discretion.
There is a softer but effective route too: practitioners setting their own rules. Academia has two well-known declarations. DORA opposes using a single metric like the Journal Impact Factor to evaluate individuals, and the Leiden Manifesto's first principle is that quantitative evaluation should support, not replace, qualitative expert judgement. These are not regulation imposed from outside, they are people inside the field agreeing that this is not how we play. Such self-regulation is often faster than waiting for regulation to land, and better fitted to the profession.
Above that comes regulation. It has real teeth: the FTC's 2024 rule comprehensively bans buying and selling fake reviews, with penalties up to 51,744 dollars per violation. But regulation is destined to arrive late: it is slow, expensive, and legislates only after the mess has grown big. That fake-review rule took years to land. So regulation is the floor, not the first choice.
Stack the four layers and it is clear why institutional defence is last and most expensive. Technical measures change code and criteria, which is fast and controllable. Institutional defence has to change incentives and power relations: who pays whom, who may challenge whom, who is accountable when things fail. Those move slowly, need coordination across parties, and usually disturb vested interests. It is expensive, but when every technical measure has been routed around, it is all that is left underneath.
To judge whether an evaluation system is institutionally healthy, ask these:
- Independence: where does the evaluator's money come from? Can it be funded by what it evaluates?
- Transparency: are the weights and methods published? Would an equally reasonable alternative algorithm shake the rankings violently?
- Appeals: can the evaluated party challenge and correct?
- Ecosystem: is one player dominant, or do several yardsticks check each other?
One open question: institutional defence cannot escape Goodhart either. An institution that exists to govern evaluation's defects has a livelihood that depends on those defects continuing to exist. So it may, without noticing, keep the problem alive (the Shirky principle mentioned in G02). So who audits the auditors? Adding another layer of auditors above the auditors just pushes the same question up one level. This infinite regress has no clean solution today.
The one-line takeaway: the ruler can be exploited no matter how you fix it, so what catches the fall is the rules about who holds the ruler and who pays them, and that layer is the slowest, most expensive and hardest to skip.
Sources / further reading
- World Bank Doing Business data manipulation and discontinuation: WilmerHale independent investigation (2021); World Bank discontinuation statement (2021-09-16); B-READY as successor (from 2024).
- The Conversation (2025-04-23), "rebrand not a revamp"; the Hong Kong SAR government challenging its B-READY rating (2025-12-30, reported by SCMP).
- The US News withdrawal wave and methodological rebuild: four tiers for medical schools (from 2024); Yale Law School falling out of the top spot for the first time in 2026 (Law.com 2026-04-06); NYT (2024-01-06) "U.S. News Makes Money From Some of Its Biggest Critics" (paid badge conflict of interest).
- Chirikov, I. (2021) (QS consulting conflict of interest: 22 of 28 Russian universities, roughly 2.85 million dollars, an average rise of about 140 places), Higher Education.
- The FrontierMath funding and exclusive access controversy — TechCrunch 2025-01-19.
- DORA (2012, sfdora.org); Hicks, Wouters, Waltman, van Eck & Rijcke (2015). "The Leiden Manifesto for research metrics." Nature 520:429–431; Wilsdon et al. (2015). The Metric Tide.
- FTC (2024). Rule on the Use of Consumer Reviews and Testimonials (bans buying and selling fake reviews, up to 51,744 dollars per violation).
- Primary sources
research/04§2A (Doing Business),research/deep/D3sub-threads 1/2 (the aftermath of withdrawal, B-READY, the badge conflict of interest), andresearch/deep/D7(institutional-layer defence and reflexivity).