REACTOR
B02 Laws Laws · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ B02 ]…

Campbell's Law

The man who warned about metrics was measurement's own champion.

Goodhart's law says a measure loses accuracy once it is treated as a target (B01). Campbell's law goes one step harder: the measure does not just lose accuracy, it drags down the thing being measured. Test scores becoming meaningless is the small loss. Education itself being remade around test prep is the big one.

Donald Campbell was a leading methodologist of the question "do social programmes actually work." His law says one thing more than Goodhart's, and it stresses that the corrosion is inevitable:

"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."Campbell, 1976/1979

In plain terms: the more a quantitative indicator is used to make important decisions, the more corruption pressure it comes under, and the more it will distort and corrupt the very process it was meant to monitor. Note the second half. The casualty is not only the number but the thing behind the number. The sentence comes from a subsection literally titled "The corrupting effect of quantitative indicators," and Campbell's own phrase is plural, pessimistic laws. He offered a set of pessimistic laws; posterity remembered only this one.

FIG.01 Campbell's tension: he wrote the corruption law and argued for the experimenting society EXPLORABLE

The examples in Campbell's 1979 paper all smell of case files. Start with the police station.

Police forces were graded on the clearance rate, the share of reported crimes that get solved. That measure produced corruption at both ends. At the reporting end, cases went unrecorded, were recorded late, or were downgraded to a lighter category. Fewer reports, smaller denominator. Campbell named Nixon's crackdown, saying its main effect was to corrupt the crime rate statistics (research by Seidman and Couzens, 1972). The solving end is more absurd: in plea bargaining, an arrested burglar confesses to several extra unsolved cases in exchange for a lighter sentence, and in the phrase Campbell quotes, "he is helping the police improve their clearance rate." Skolnick's 1966 research points out that many of them confessed to crimes they had not committed. The better the clearance rate looks, the further the files drift from the truth.

Now the schools. The Texarkana performance contracting case (documented by Stake, 1971): an education contractor was paid according to how much student scores rose, so it simply taught the final exam questions. Scores rose, the money was paid, and the students learned nothing extra. Then comes the passage of Campbell's that gets quoted most: achievement tests may well be valuable indicators of general school achievement under conditions of normal teaching aimed at general competence, but when test scores become the goal of the teaching process, they both lose their value as indicators of educational status and distort the educational process in undesirable ways.

His evidence also crosses borders and battlefields: the Soviet output-target literature (Granick 1954, Berliner 1957), and enemy kill counts in Vietnam, that is, body count (grading war performance by number of corpses reported). Once kill count became a hard metric for grading units, units came under pressure to inflate the number, even to kill indiscriminately to pad it. The My Lai massacre, in his account, is exactly that pressure pushed to its extreme, in battlefield form. That thread opens out in B05 (the McNamara fallacy).

The easiest way to misread this law is as a manifesto against quantification. Campbell's academic identity was the exact opposite. The textbook on experimental and quasi-experimental design he co-wrote with Stanley is the founding text of the entire programme evaluation field, and his 1969 paper "Reforms as Experiments" called for treating social reform as experiment and guiding policy with rigorous measurement. He was the strongest standard-bearer of the "experimenting society."

The tension is that in the same 1979 paper he both calls for large scale quantitative evaluation and asserts that quantitative indicators will inevitably be corrupted. His reconciliation is not to abandon measurement but to design institutions: bring in external evaluators; let stakeholders such as unions act as watchdogs; triangulate across multiple indicators; keep indicators as far as possible from direct reward and punishment loops. In his own words, we should study the social processes through which corruption is exposed and try to design social systems that incorporate those features. The law is therefore a built-in safety clause inside the experimenting society, not a repudiation of it. That line of thought later grew into the whole defence branch: G04 on multi-indicator checks and balances, G06 on institutional-layer defences.

Read Campbell and Goodhart side by side and the division of labour is clean. Goodhart's subject is the statistical regularity: the correlation between measure and goal collapses under control pressure, and the loss happens at the measurement end. Campbell's subject is the social process: teaching, policing, adjudication, the measured activities themselves get reshaped by the indicator, and the loss happens at the world end. The first says the ruler goes out of true, the second says the ruler rewrites what it measures. That is also why Campbell chose "corrupt" rather than "distort." Distortion can still be calibrated away; corruption changes the organization that produces the data. This course treats that half step as the interface between the blue and red branches: blue handles the formal relation between ruler and goal, red (starting from R01, reactivity) handles what the ruler does to the world.

The counter-evidence belongs on the table too. That is the line between serious documentation and a collection of cautionary tales.

Dee and Jacob in 2011 used the timing differences in when US states introduced accountability to test the No Child Left Behind Act (NCLB, the federal law tying school funding to standardized test scores): by 2007, fourth grade mathematics had risen about 0.22 standard deviations on NAEP. NAEP is a national sampled audit test with low stakes, so teachers have no reason to teach to it, which makes that gain more likely to be real learning. What does 0.22 standard deviations mean? Roughly moving a middle-of-the-pack student to the upper middle, which counts as a fairly substantial effect for an educational intervention. And the gain shows up across all five mathematics subscales, not only in the easily coachable item types. In fairness, the same study found no effect in reading at either grade level, and it also documents the cost of test subjects crowding out other classes.

The inflation side of the picture appeared earlier: Cannell's 1987 survey found that every state claimed to be above the national average on its own high stakes state test. Statistically it is impossible for everyone to be above average, so this collective illusion got the name the "Lake Wobegon effect" (a fictional town where all the children are above average).

The combined reading is this: high stakes scores carry both a real-gain component and an inflation component. Campbell's law asserts that the second one necessarily exists, not that the first one is necessarily zero. One ungameable audit channel (a test like NAEP that nobody has an incentive to teach to) can separate the two. The full boundary conditions of the law are in B14 (conditions for benign indicators), and the complete body of evidence from the education battlefield (score inflation in Kentucky's KIRIS, the teacher answer-sheet tampering detected in 4% to 5% of Chicago classrooms) is in K01 (the education file in the case band).

One open question: Campbell's prescription relies on watchdogs and external evaluators, but the watchdog itself has to be graded, and the audit channel is itself an indicator. What stops the channel that exposes corruption from being corrupted on the second pass? This recursive problem gets half an answer each in R09 (the audit society) and G06 (institutional-layer defences).

The one-line takeaway: the blast radius of a high stakes indicator goes beyond the number; it remakes the activity that produces the number.

Sources / further reading
  • Campbell, D. T. (1979). "Assessing the impact of planned social change." Evaluation and Program Planning 2(1):67–90 (= the 1976 Occasional Paper #8; read at a conference in 1974).
  • Skolnick (1966), Seidman & Couzens (1972), Stake (1971), Granick (1954), Berliner (1957): sources of the examples cited in Campbell's original.
  • Campbell, D. T. (1969). "Reforms as Experiments." American Psychologist 24(4):409–429 (the experimenting society programme).
  • Rodamar, J. (2018). "There ought to be a law! Campbell versus Goodhart." Significance (priority review).
  • Dee, T. & Jacob, B. (2011). "The impact of No Child Left Behind on student achievement." JPAM 30(3):418–446 (counter-evidence).
  • Cannell, J. J. (1987). The Lake Wobegon report (survey of state tests where everyone is above average).
  • Documentation and counter-evidence in research/03 §3, §9 and research/deep/D2 §3.
Where to next