Test scores rise while the thing they are supposed to represent does not. You have probably lived through this. Education is where Campbell's law originated and the field where its evidence is most complete. The chain of facts comes in three parts, like a slope of rising pressure: scores on high-stakes tests can rise steadily while real ability stays flat (score inflation, with econometric evidence from cross-test comparison); when the pressure gets high enough, inflation crosses the line from legal into collective fraud (Atlanta, with criminal conviction records); and the same accountability regime also produced real gains on low-stakes audit tests (NCLB's math results). Same teachers, same accountability regime: switch to a low-stakes ruler and it measures real progress, switch to a high-stakes ruler and it measures inflation or even fraud. What changed is the ruler, not the people. Together the three parts give the conclusion: teaching to the test is not a story of moral decline among educators, it is the structural consequence of tying high stakes to a distant proxy metric. The width of the gap between scores and ability depends on metric design, not on teachers' character.
Where Campbell's law originated
Start at the first scene. You read Campbell's law in B02: the more social decision pressure a metric carries, the more it is liable to corrupt, and the more it corrupts the process it was meant to measure. When Campbell put the proposition forward in the 1970s (publication details in the sources), education was the first illustration: in the Texarkana performance contracting case (documented by Stake 1971), contractors paid according to gains in student scores taught the final exam items directly. With money tied to gains, the cheapest source of gains is the items themselves. Campbell's own words on testing are worth reading in full:
"Achievement tests may well be valuable indicators of general school achievement under conditions of normal teaching aimed at general competence. But when test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways."Campbell, 1976/1979
In plain terms: under normal teaching, tests are decent indicators; but once scores become the goal of teaching, two losses happen at once. The indicator fails (scores stop indicating ability) and the process corrupts (teaching itself is rebuilt around the test). Fifty years of empirical work since is the unfolding of those two sentences. Mechanism tag: adversarial (the tagging system of K01, explained in detail in B08), with the levers pushed to maximum being stakes attachment and the manipulable channel.
The econometrics of score inflation
Score inflation is the central concept of the measurement scholar Koretz: gains on high-stakes tests far exceed the real growth in ability shown by low-stakes audit tests. The cleanest case is in Kentucky: from 1992 to 1995, scores on the state test KIRIS, the target of accountability, rose sharply, while the same cohort's ACT scores (a college entrance test nobody was preparing for) did not budge. The gap was about 0.7 standard deviations in math and about 0.4 in reading (Koretz & Barron 1998, a RAND report). What is 0.7 standard deviations? Roughly the gain that would lift a middling student into the top quarter, all of it on paper. Earlier there was the "Lake Wobegon effect": Cannell found in 1987 that every state claimed its scores were above the national average, which is arithmetically impossible, showing that inflation was already a nationwide phenomenon.
The curve Koretz displays repeatedly is therefore called the sawtooth: a new test comes into use, scores first fall back to the real level, then climb steeply year on year; switch to the next test and they fall again, then climb again. What the rising segment measures is not ability growing but teachers' and students' familiarity with that particular item format growing. Each falling edge is a natural audit, exposing the inflated portion of the previous climb. This makes the audit channel the key instrument in this field: low-stakes tests like NAEP that nobody prepares for can separate real gains from inflation. Mechanism tag: rules arbitrage (test-format-specific specialization), not fraud; it is entirely legal, which is exactly why it is far more common than cheating.
Teaching up to the pass mark and no further
How an accountability regime counts determines the shape of the distortion. NCLB (the US federal education accountability law from 2002) took "number of students meeting the standard" as its core yardstick: only whether you cross the line counts, and how far past or short of it you land adds nothing further. So the rational response for teachers is to direct resources at students sitting just around the pass mark (bubble kids): Neal & Schanzenbach 2010 used Chicago data to show that mid-range students' scores were lifted while both the weakest and the strongest were neglected at the same time, because neither end makes any marginal contribution to the metric. Abandoning them is arithmetic, not cruelty. Grading is reactive too: New York State Regents exams were graded by teachers at the same school, and the score distribution piled up abnormally just above the pass mark. Dee et al. estimated in 2019 that about 40% of scores that should have fallen just below the pass mark were manually lifted above it, meaning nearly half of the near-misses were rescued by the grader's pen. Total cheating has an algorithmic estimate too: Jacob & Levitt 2003 used a detection pattern combining anomalous score jumps with suspicious answer strings and estimated that in Chicago at least 4% to 5% of classrooms each year had answer sheets altered by teachers or administrators, with the proportion rising with accountability pressure. The three pieces of evidence compose one distribution: legal test-format specialization is the bulk, lenient grading is the middle band, and outright fraud is the tail. The greater the pressure, the fatter the tail.
Atlanta: the judicial record of the tail
The Atlanta Public Schools answer-changing case is the most complete record of that tail to date. The driver was NCLB's Adequate Yearly Progress targets (AYP) and the bonuses and job security attached to them, with pressure pushed down layer by layer from the superintendent. From 2009 the Atlanta Journal-Constitution used erasure analysis (statistics on abnormally dense wrong-to-right erasures on answer sheets) to question the surging scores; in July 2011 a Georgia Bureau of Investigation report found cheating at 44 schools with about 178 educators implicated, using methods including collectively altering answer sheets after the exam, and even holding group "erasure parties". In March 2013 a grand jury indicted 35 educators on charges including RICO, a statute originally intended for fighting the mafia and organized crime; most defendants pleaded guilty, and in April 2015, 11 of the 12 who went to trial were found guilty under RICO, with some initial sentences of up to roughly 20 years (later reduced through sentencing negotiations). Superintendent Beverly Hall, once named national superintendent of the year, died of illness in March 2015 before trial and was never tried. Prosecuting teachers under a statute for fighting the mafia shows the justice system concluded this was not individual lapses but an organized network of fraud. Mechanism tag: outright fraud. It and test-format specialization are not two phenomena but the two ends of one pressure gradient: how tightly bonuses and job security are attached to the metric determines where along the spectrum behavior lands.
The national examination hall: the PISA shock
Scale the examination hall up to the national level and the mechanism is unchanged. PISA is the OECD's cross-national test of student competence, effectively a common examination that ranks countries. The first round in 2000 came in below the OECD average and set off the "PISA-Schock" in Germany, which held the headlines for weeks, after which the state education ministers introduced national education standards and an all-day schooling programme. Japan's reading rank fell from 8th in 2000 to 15th in 2006, and public opinion put the blame on "relaxed education", which had cut roughly 30% of curriculum content. Curriculum revisions in 2008 and 2011 added teaching hours back, effectively ending relaxed education. One cross-national ranking rewrote the timetables of two major education systems: that is the force of reactivity at national scale. The other side of Shanghai's two first-place finishes in 2009 and 2012 is the sample: because of the household registration barrier, over 120,000 fifteen-year-old children of migrants were outside the school system being sampled, biasing the result upward (Loveless's critique), and representativeness was questioned again after the 2018 switch to joint sampling across four provinces and cities. Mechanism tag: policy convergence compounded by a selection effect at the sample level. Countries do what districts do: change the yardstick they can change first.
The counter-case: when teaching to the test is learning
Telling education accountability as a pure Campbell disaster is equally unfaithful to the evidence. Dee & Jacob 2011 used the staggered timing of state accountability adoption for identification and estimated NCLB's causal effect: by 2007, fourth-grade math on NAEP had improved by about 0.22 standard deviations, and the gains were spread across all five math subscales rather than concentrated in easily coached item formats. NAEP is an audit test nobody prepares for, which makes real learning gains the more likely explanation. 0.22 standard deviations is not small, roughly equivalent to a few months of extra learning, and it appeared on a ruler nobody was aiming at, which is the strongest evidence that accountability really taught something. The costs are in the data as well: no effect in reading at either grade, and tested subjects crowding out other classes. Koretz and Dee/Jacob do not contradict each other: high-stakes scores contain both real gains and an inflation component at once, and audit tests can separate the two. The conditions under which teaching to the test converges on learning are the education-side projection of B14 conditions for a benign metric: what is tested is close enough to the ability you want (short proxy distance), an ungameable audit channel exists, and the stakes are not strong enough to push people past the legal line.
One open question: value-added accountability, which replaces the count meeting the standard with growth, should in theory weaken the bubble kids effect, because every student's marginal progress counts. But the empirical evidence since Race to the Top is not sufficient. Was the bubble eliminated, or did it just move somewhere else?
The one-line takeaway: rising scores say nothing about learning improving, unless you also have a second ruler that nobody is aiming at.
Sources / further reading
- Campbell, D. T. (1979). "Assessing the Impact of Planned Social Change." Evaluation and Program Planning 2(1):67–90. Publication history: the proposition was read at the Visegrád conference in Hungary in 1974, circulated in 1976 as Western Michigan University Occasional Paper #8, and formally published in 1979, with the section heading "The corrupting effect of quantitative indicators".
- Koretz, D. & Barron, S. (1998). The Validity of Gains in Scores on the KIRIS. RAND MR-1014; Koretz (2008) Measuring Up; Koretz (2017) The Testing Charade.
- Jacob, B. & Levitt, S. (2003). "Rotten Apples." QJE 118(3):843–877.
- Neal, D. & Schanzenbach, D. W. (2010). "Left Behind by Design." Review of Economics and Statistics.
- Dee, T., Dobbie, W., Jacob, B. & Rockoff, J. (2019). "The Causes and Consequences of Test Score Manipulation." AEJ: Applied Economics.
- Dee, T. & Jacob, B. (2011). "The Impact of No Child Left Behind on Student Achievement." JPAM 30(3):418–446.
- Atlanta case timeline: GBI investigation report (2011), CNN (2013-03-29), NYT (2015-04-01).
- The PISA shock and the Shanghai sampling controversy:
research/04-cases-across-domains.md§2B (including the critique by Loveless, Brookings). - Working notes:
research/03-goodhart-family.md§6.4, §9.1;research/deep/D3§4.2–4.3.