REACTOR
K02 Cases Cases · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ K02 ]…

Governance: Targets & Terror

The four-hour miracle on the waiting list.

Requires R04 Discipline & Surveillance Unlocks

Attaching a metric to someone's job is the most common move governments use to raise the efficiency of governance, and it is where reactivity runs hottest. Three cases occupy three segments of the mechanism spectrum: England's public hospitals were governed by a four-hour target, and doctors learned to stop the clock, but waiting really did get shorter (rules arbitrage, with control evidence of genuine improvement); the New York Police Department was held to account on crime numbers week by week, and felonies quietly got recorded as misdemeanours (classification manipulation, with the underlying decline real); the World Bank ranked countries on their business environment, and eventually a large ranked country reached its hand into the ranking itself (measurement capture, with the ranking body itself drawn into the game). The shared skeleton of all three is the disciplinary structure described in R04: publish on a schedule, hold each level accountable, and remove those who miss the target. Every layer then has ample motive to make the numbers look good, and which method gets used depends on which channel is cheapest.

The NHS: targets and terror

In the 2000s England fitted its hospital system with star ratings and executive accountability: miss a target such as waiting time and managers lost their jobs. The researchers Bevan and Hood gave this regime a name in 2006, targets and terror, and stated in their opening that it bore an obvious resemblance to the Soviet command economy. Their theoretical term is synecdoche, the rhetorical figure of the part standing for the whole: any target regime can only extract a few measurable dimensions to represent overall performance. Managers' attention and resources then follow those same measured dimensions, because only those get seen and held accountable. So the dimensions that were not chosen become a dumping ground. Take the four-hour A&E target (a patient must be seen or admitted within four hours of arriving at A&E). The working notes establish the following list of methods.

  • Ambulance stacking: patients held in ambulances outside the A&E department, where the four-hour clock does not start (with recorded fatal cases of delayed critical care).
  • Surges during measurement weeks: extra overtime staff drafted in and elective surgery cancelled during the week under assessment, with the other fifty-one weeks reverting to normal.
  • Short-stay observation beds: patients approaching the limit moved into observation beds, which under the metric's definition no longer count as A&E waiting. Later research confirmed that these clock-stopping methods were widespread behavior, amounting to a redefinition of admission.
  • Retrospective record correction: in 2003 the Commission for Health Improvement found that about a third of ambulance trusts had "corrected" response time data, with the reported distribution showing an anomalous spike at the 8-minute target. A third is not a few bad apples, it is the scale of an industry norm.
  • The gap between the yardstick and the experience: in 2002/03, 139 of 158 acute trusts reported that 90% of patients were seen within four hours, while only 69% of patients reported that experience in surveys. Ninety percent on paper, seventy percent in experience, and the crack between them is the volume of definitional manipulation.

Theoretically Bevan and Hood distinguish three kinds of distortion: the ratchet effect (this year's target becomes next year's baseline, so nobody dares overshoot much), the threshold effect (everyone converges on the minimum passing line), and output distortion (the numbers improve while the real service does not change). Mechanism tag: a full-spectrum demonstration of rules arbitrage and definitional arbitrage, with the yardstick in the hands of those being assessed.

But the NHS is also the best teaching material for "reactivity does not equal pure fraud", because it has a clean control group: Scotland is also part of the NHS and did not adopt England's target accountability regime. Propper et al.'s difference-in-differences in 2008 (comparing the change in each place before and after and subtracting the common trend) shows that England's proportion of long waits fell significantly relative to Scotland, with no clear evidence found of widespread distortion of clinical priority; waiting times also fell in the lower ranges that were not targeted, which further weakens the pure-gaming explanation. In other words, after subtracting every clock-stopping trick, a block of genuine improvement remains. Bevan and Hood themselves wrote in the companion BMJ piece that nobody wants to return to the pre-target NHS: back then over 20% of patients waited more than four hours in A&E, and elective admission took 18 months. The IFS assessment in 2020 points the same way: the four-hour target did shorten waits by about 20 minutes and may have saved lives, at the cost of admitting marginal patients to stop the clock, which pushed up admission rates and costs. Twenty years on, the gaming is still evolving: in early 2026 British media and audit bodies questioned the government paying per head (about 33 pounds per patient removed) to flatter waiting lists. This is still developing and should be treated as unsettled evidence.

CompStat: a decline curve produced by downgrading

CompStat was introduced at the NYPD in 1994 by Commissioner Bratton and Maple: crime data precise to the precinct, with weekly accountability meetings pressing precinct commanders. It was long treated as the model of data-driven governance, until insiders spoke. The main figures exposing manipulation were the criminologist Eterno (himself a retired captain) and Silverman, who went and asked retired officers directly. In the first survey round, 157 of 309 respondents knew of crime reports having been altered after the fact, and of the 160 who answered the ethics question, only 22.5% considered the alterations ethical; in a 2012 expanded survey of 1962 retired officers of all ranks, over half of the 871 who retired after 2002 reported personal knowledge of manipulation. A knowledge rate above half means this was not deviance at a few precincts but routine institutional practice. The methods are structurally identical to the NHS, with the yardstick swapped for offence classification: downgrading grand larceny to petty larceny, lowballing the value of stolen property to change the charge, refusing to take reports, and declaring reported cases unfounded. Independent corroboration comes from Schoolcraft's precinct recordings (surfaced in 2010) and reviews by the department's own quality assurance division. The causal evidence is Eckhouse's 2022 cross-city study: adopting the CompStat system is associated with roughly 3500 additional minor arrests per city per year, while felonies did not actually decline. Metric management produced arrest volume, not public safety.

The balance has to be given its full weight too: New York's large crime decline since the 1990s is substantially real, and categories such as homicide, which is nearly impossible to downgrade, declined just as much (the conclusion of Zimring and others). A body cannot be reclassified as a graze, so the homicide curve is the hardest control ruler to fake. The accurate statement is that manipulation concentrated on the manipulable yardstick, namely the boundary between felony and misdemeanour, not on all metrics. The technique is not a New York specialty either: in 2025 a commander at the Washington DC police department was suspended and investigated for shaving felonies down to misdemeanours. Mechanism tag: classification manipulation plus goal substitution; Campbell's 1979 prediction about clearance rates (B02) replaying unchanged in the age of data-driven management.

Doing Business: capturing the measuring device

The World Bank's Doing Business report scored and ranked the regulatory environment of about 190 economies, directly shaping national reform agendas: India's prime minister Modi publicly demanded that the country's rank rise 100 places, and several countries set up dedicated ranking task forces. When the consequences are large enough, the gaming no longer stays on the side of the measured. The independent investigation by the law firm WilmerHale, commissioned by the World Bank's ethics committee (reviewing over 80,000 documents), reconstructed the core events around the 2018 edition: the draft placed China 85th; Chinese officials, representing the Bank's third-largest shareholder, applied repeated pressure; then-CEO Georgieva took charge and instructed the team to find "adjustable" data points, and changes in three areas of ambiguous judgement eventually moved China up 7 places to 78th, exactly level with the previous year. The signature detail in the investigation is that Georgieva went personally to a project manager's home to collect the paper final draft reflecting the changes and to thank him: going to someone's house to collect the working copy shows the people involved knew these changes could not stand daylight. The backdrop was a 13 billion dollar capital increase negotiation, raising China's shareholding from 4.68% to 6.01%. Then chief economist Romer had earlier (January 2018) said publicly that the indicators were contaminated by political motives, and resigned 12 days later (later partly retracting his wording). In August 2020 the Bank suspended publication and launched an audit, and on 16 September 2021 announced that the whole ranking was permanently discontinued. A global ranking that ran for nearly two decades died of its own leadership changing the numbers: the mechanism tag is measurement capture, the far end of the spectrum, where the ranked party and senior figures at the ranking body collude to alter the measurement itself.

Discontinuation was not the end of the story. The successor index B-READY launched in 2024 (covering about 50 economies) and expanded to about 101 in 2025, with a methodology claimed to be more transparent. After the second report was released in December 2025, the Hong Kong SAR government immediately issued a public statement that the rating was outdated and unfair, on the grounds that local data was collected in 2023 while other economies used 2024 data. From secret capture back to open protest, the intensity of reactivity dropped a notch while the motive structure stayed exactly as it was: as long as a ranking has consequences, the ranked will keep contesting the scoring itself.

One open question: how many percentage points did CompStat-style downgrading contribute to New York's crime decline curve? Eckhouse gives a cross-city average, but there is still no clean estimate of the magnitude of downgrading in New York specifically. In mixed cases where the bulk is real and the manipulation local, quantifying the manipulated share precisely is the most unresolved technical task in this field.

The one-line takeaway: pressure does not select for character, it selects for channels; whoever controls the yardstick is where the numbers deform, and when the pressure is high enough even the people setting the test join in.

Sources / further reading
  • Bevan, G. & Hood, C. (2006). "What's Measured Is What Matters: Targets and Gaming in the English Public Health Care System." Public Administration 84(3):517–538; companion BMJ piece 332:419–422.
  • Propper, C., Sutton, M., Whitnall, C. & Windmeijer, F. (2008). "Did 'Targets and Terror' Reduce Waiting Times in England for Hospital Care?" B.E. Journal of Economic Analysis & Policy 8(2).
  • IFS (2020). Assessment of the four-hour target; the waiting-list "payment per head" controversy in The Times / NAO (2026-02, developing).
  • Eterno, J. & Silverman, E. (2012). The Crime Numbers Game: Management by Manipulation. CRC Press.
  • Eckhouse, L. (2022). "Metrics Management and Bureaucratic Accountability: Evidence from Policing." AJPS.
  • WilmerHale (2021). Doing Business independent investigation report; World Bank discontinuation statement (2021-09-16); the Hong Kong government's statement on B-READY (2025-12-30).
  • Working notes: research/03-goodhart-family.md §6.6–6.7, §9.2; research/04-cases-across-domains.md §2A; research/deep/D3 sub-threads 2 and 4.
Where to next