REACTOR
B14 Laws Laws · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ B14 ]…

When Metrics Work

When is teaching to the test just teaching? The law has edges.

Requires B10 Baker Distortion Unlocks

Everything so far has been about how measures go bad. Listing only failure cases is itself dangerous: this list itself was picked out by the result "the measure failed", and picking evidence by the result is exactly the flaw this branch teaches you to spot, selection bias. So there has to be a case for the defence, stated with the same pedantry used on the failures, about when measures really do work and which hard evidence stands on their side. The point is not to split the difference, it is to correct Goodhart's law from a curse into a conditional proposition: it has conditions under which it holds, and outside those conditions measurement remains the strongest management technology known to us.

The first weapon for the defence you already met in B10, only there it was used to convict and here to acquit. Baker's criterion says a measure's quality turns on directional alignment ρ(Pₑ, Vₑ): when it approaches one, each of your actions affects the measure in exactly the same direction and proportion as it affects true value. Same direction, same proportion means fully optimizing the measure is fully optimizing value; Firm A's sales commission, therefore, can be maxed out with confidence. Translated to the exam hall: when what is tested is the capability you want, and raising your score through real skill is cheaper than raising it through tricks, then studying for the test is learning. Grinding algorithm problems raises exactly your ability to write algorithms, memorizing TOEFL words raises exactly your vocabulary, and Goodhart has no part in those scenes. Two echoes on the theory side: read El-Mhamdi and Hoang's tail-distribution criterion (from B08) in reverse, and when the difference distribution is not heavy tailed, over-optimization does not damage the true goal; the 2024 cross-disciplinary proxy failure target article likewise stresses that proxy failure is often only partial failure, not all or nothing.

The hardest empirical evidence for the defence comes from British healthcare. Set the background straight first: in the 2000s the English NHS (National Health Service) imposed hard waiting time targets on hospitals, four hours in accident and emergency and limited waits for elective surgery, with management at failing hospitals liable to be removed or publicly named and shamed. Academics nicknamed the regime targets and terror. The critics Bevan and Hood documented plenty of gaming: ambulances held outside A&E without unloading patients, crash overtime to get through a measurement week (a whole lesson on that dispute is in K02). Following the logic of the first half of this branch, you would expect the regime to have been a pure disaster.

But the data tells the other half of the story. The economist Propper and colleagues seized a natural control in 2008: Scotland has a structurally identical NHS but did not adopt England's target regime. The difference-in-differences result (comparing the change over time in one place against the other) is that the target regime did significantly cut waits in England, and no clear evidence can be found that this was bought through falsified data or sacrificed clinical quality. One further detail is devastating: waiting times also fell in the lower ranges that the targets did not aim at. If the improvement had been pure gaming, the numbers should have piled up near the target line only, and the whole distribution should not have improved. Even the critics concede: in their companion piece in the BMJ, Bevan and Hood wrote that nobody wants to go back to the NHS as it was before targets, when more than a fifth of patients waited over four hours in A&E and elective admission meant waiting eighteen months or more. Why the measure won in this case: waiting time is not a proxy for patient welfare, it is part of patient welfare. There is almost no distance between ruler and goal, so there is little room to game.

Education accountability gives a subtler two-sided sample, and both sets of numbers are worth remembering. The positive side: the US No Child Left Behind Act, effective from 2002, mandated national accountability for schools on test scores. Dee and Jacob in 2011 used the timing differences in when states adopted accountability for identification and found NCLB raised fourth grade mathematics by about 0.22 standard deviations by 2007. What matters is which ruler that number sits on: NAEP, a low stakes national sampled audit test. It affects no school's funding or ranking, so teachers have no reason to coach for it. And the gain spans all five mathematics subscales including algebra and measurement, rather than concentrating in the item types that are easy to predict. Together those two points point to real learning rather than score inflation. The cost is recorded honestly too: reading showed no movement at either grade level.

The negative side: the classic audit by the education measurement scholar Koretz. Kentucky's high stakes state test KIRIS produced large score gains from 1992 to 1995, yet when the same students sat the low stakes ACT the results did not budge, with a gap between the two rulers of about 0.7 standard deviations in mathematics and about 0.4 in reading. The part of achievement that the state test showed evaporated when the ruler changed, and that is score inflation. The two bodies of evidence do not conflict, they assemble the full picture: high stakes score = real gain + inflation component, and a low stakes audit test is the filter paper that separates the two. A test score is a distal proxy for achievement (long proxy distance), so distortion is heavy, which again throws into relief how lucky the NHS case was in having "the ruler is the goal" (the education dispute opens out in K01).

Step outside assessment settings and the evidence comparing "with measures" against "without measures" is equally one-sided. After Los Angeles required restaurants to post hygiene grade cards in the window in 1998, Jin and Leslie measured a triple hit: kitchen hygiene inspection scores rose, customers began voting with their feet, and hospital admissions for foodborne illness fell. The measure did not stay a ranking game, it changed the number of bacteria in the real world. Hastings and Weinstein's experiment found that putting school performance information directly into the hands of low income families significantly raised the share of parents choosing high performing schools, with children's results benefiting accordingly. The mechanism behind this: a public measure turns information formerly available only to the well connected into a public good, and the distributional effect therefore favours the weak. The historian of science Porter's famous claim is exactly this: quantification is most valuable precisely where trust is scarce and elites do not get the last word, as a weapon outsiders use to constrain unaccountable discretion. Quoting only his "numbers are cold" without this half is quoting out of context.

The route of "skip measures and rely on human judgment" has been tested by psychology for half a century. Grove and colleagues' meta-analysis of 136 studies across fields found that expert informal judgment beats a simple statistical formula in fewer than 5% of cases; in Schmidt and Hunter's century-long synthesis, a structured interview with a scoring rubric has predictive validity around 0.51 while a free-form unstructured interview only about 0.38. Unstructured human judgment is full of noise, bias and cronyism, and quantification and structure are often both more accurate and fairer. Same on the management side: Bloom and Van Reenen's cross-country survey found that performance monitoring practices correlate robustly with productivity, and a follow-up randomized experiment in Indian textile mills supplied the causal evidence, with plants that introduced structured monitoring and targets raising productivity by about 17%. Lazear's study of a windscreen fitting company shows that when the task is single-dimensional and output verifiable, piece rates are simply clean and effective (output up, no quality disaster). This is the mirror image of the theorem in B09: when a task has only one dimension and quality is verifiable, multitask crowding out cannot occur.

Finally, back to evals themselves. Theory predicted long ago that repeatedly selecting models on the same test set slowly leaks the test set into your selection, so scores must inflate (a whole lesson on that fear is in Y02). The empirical verdict is unexpected. Kaggle is the largest machine learning competition platform, and every contest comes with a natural control: the public leaderboard (which entrants can hammer repeatedly) and the private leaderboard (held-out data revealed only at the end). Roelofs et al. reviewed over a hundred competitions in 2019, and their conclusion reads: somewhat surprisingly, little evidence of substantial overfitting. The gap between public and private boards is mostly explained by random variation, and it holds across domains, loss functions and entrant populations. Mania and colleagues supplied the mechanism: competing models are highly similar to each other, so no matter how many times you submit, you are really asking the same repeated question, and the effective number of independent queries against the test set is therefore far below the number of submissions. Fewer effective queries means the overfitting budget is spent more slowly, slower than the theoretical bound. Reviews on question answering benchmarks are equally clean: nearly all the score drop comes from new questions being slightly harder, not from the leaderboard being worn out. Record the lesson precisely: theory guarantees degradation in the worst case, while experience shows that degradation in the average case is diluted by model similarity to the point of being nearly unmeasurable. Confusing the two turns "evals get faker the more they are hammered" into an unconditional law. Record the boundary too: this verdict comes from supervised learning competitions before 2019, and the adaptive manoeuvres of the LLM era (tuning prompts, changing scaffolds, post-training) are an entirely different shape with no review of comparable scale yet.

Put Muller, Koretz, Bevan and Hood together with the positive evidence and you get five operational criteria for a measure trending benign, each checkable. One, short proxy distance: what is measured is (or is very close to) what you actually want. Two, an ungameable audit channel exists: keep one ruler tied to no reward or punishment, purely for verification. Three, raising real output is cheaper than gaming: especially true when the task is one dimensional and output verifiable. Four, low or medium stakes use: the measure serves diagnosis and learning rather than mechanically triggering reward and punishment. Five, multiple measures with human judgment as a backstop, rotated periodically to counter adaptation. All five holding at once is rare, but the more that hold, the closer the measure comes to its ideal identity: a measuring instrument rather than a target in a game. One sentence to close the blue branch: Goodhart's law is not a claim that measurement is useless, it is the failure law of one combination, proxy measure plus high pressure optimization. Remove the pressure, shorten the proxy chain, keep the audit and the judgment, and measurement remains the strongest management technology known. And do not forget that Campbell himself was a standard-bearer of quantitative evaluation, and his pessimistic law is a built-in safety clause inside the "experimenting society," not a repudiation of it.

One open question: the defence's prettiest evidence (the hundred-competition Kaggle verdict) comes from old-style competitions where you submitted a prediction file, and today's adaptivity of tuning prompts, swapping scaffolds and post-training against a leaderboard is far more aggressive than back then. Can model similarity still protect the holdout, or will the LLM era eventually deliver the large scale adaptive overfitting that theory predicted long ago? No review of comparable scale currently answers that.

The one-line takeaway: the measure itself is not the poison, the cocktail of "distal proxy plus high stakes plus no audit" is; get the conditions right and studying for the test is learning.

Sources / further reading
  • Propper, C., Sutton, M., Whitnall, C. & Windmeijer, F. (2008). "Did 'Targets and Terror' Reduce Waiting Times in England for Hospital Care?" B.E. J. of Economic Analysis and Policy 8(2); Bevan, G. & Hood, C. (2006). Public Administration 84(3) and BMJ 332:419–422.
  • Dee, T. & Jacob, B. (2011). JPAM 30(3):418–446 (NAEP +0.22 SD); Hanushek & Raymond (2005). JPAM 24(2); Koretz & Barron (1998). RAND MR-1014 (KIRIS against ACT); Koretz (2017). The Testing Charade.
  • Jin & Leslie (2003). QJE 118(2); Hastings & Weinstein (2008). QJE 123(4); Porter (1995). Trust in Numbers; Grove et al. (2000). Psychological Assessment 12(1); Schmidt & Hunter (1998). Psychological Bulletin 124(2); Bloom & Van Reenen (2007). QJE 122(4) and Bloom et al. (2013). QJE 128(1); Lazear (2000). AER (Safelite).
  • Roelofs et al. (2019). "A Meta-Analysis of Overfitting in Machine Learning." NeurIPS (the hundred Kaggle competitions); Mania et al. (2019). NeurIPS (model similarity); Miller et al. (2020). ICML.
  • Muller (2018). The Tyranny of Metrics (base text for the criteria list); El-Mhamdi & Hoang (2024); John et al. (2024). BBS 47:e67 (partial failure).
  • The full case for the defence in research/06 part B and research/03 §9–10; cases in research/04; deepening in research/deep/D2 §5, §7 and research/deep/D4 §1.
Where to next