The lowest-maintenance rule there is: never let an eval score directly decide whether to change something, or how. Once a score carries stakes, the thing being measured learns to please the score instead of doing the job well. Separate measurement from reward and punishment. Let the eval be a health check, not a machine that hands out prizes automatically. That move is called decoupling.
You saw requisite variety in B11: a ruler that is too simple cannot govern a complex system. The audit society in R09 made a related point: once "being auditable" becomes the goal in itself, things start to warp. Decoupling turns both into one thing you can do today.
Start with an old experiment. Deming, the founding figure of quality management, gave workers a paddle with holes in it and had them scoop beads from a bucket of mixed red and white beads. Red beads counted as defects. How many red beads came up was pure luck on that scoop, and no amount of care changed it. Management nevertheless ranked people every day, best and worst, praising some and dressing down others. Deming's point was blunt: in a place where outcomes are determined by the system and by luck, ranking and rewarding individuals is rewarding and punishing luck as if it were skill. Unfair and useless.
He also had the funnel experiment, which covers the other half. Drop beads through a funnel at a target. If you correct the funnel each time by however much the last drop missed, the scatter gets worse. Leave the funnel alone and you get the tightest grouping. Taking each small wobble in an eval score and tuning parameters off it immediately is exactly this kind of over-correction. You are treating noise as signal.
"Scores get pleased once they carry stakes" is not an abstraction. Two ready examples. The UK NHS once set a hard four-hour target for A&E handling and tied it to performance reviews. Some hospitals responded by holding patients in ambulances and not letting them through the door, because the clock started at the door. On paper nothing exceeded four hours, and patients were not treated any faster. It is the same in AI. On some leaderboards vendors privately test many versions and post only the highest-scoring one, and models gradually learn to flatter whoever is grading, writing longer and more agreeable answers. Scores went up, capability did not. The common thread: as long as there are real consequences behind a score, the party being measured spends its effort pleasing the score.
A harder point: why a single metric cannot govern a general system is not a matter of "the metric was designed badly". It is mathematically impossible. Cybernetics has a law saying that how many kinds of disturbance a regulator can suppress depends on how much variety the regulator itself has. A score has one degree of freedom, high or low. Very little variety. A general system can vary across thousands of dimensions. Enormous variety. Compare the two, and you have really only pinned down the one dimension you measured. The system drifts freely in every dimension you did not measure. And "cater to what is measured, sacrifice what is not" is precisely the definition of gaming. To actually govern it, either build the ruler as a set that measures several directions at once (that is the job of G04), or accept that the system will drift where you are not looking. There is no third option.
Twist all of that into one prescription and you get decoupling: use the eval only to watch trends, find faults and raise questions, never directly to decide rewards, punishments or the direction of changes. Carter has a good image. Most indicators are really tin openers, whose job is to pry open a can of problems and force you to ask "why is this number so strange", not dials whose reading turns the steering wheel. The moment you use an indicator as a dial and wire consequences to it, the path from number to reward becomes a mechanical link, and that mechanical link is exactly what makes gaming profitable. Change it to "an anomalous number triggers a human investigation" and the expected return on manipulating any single number drops immediately.
But decoupling is not set-and-forget. It backslides on its own. The UK NHS post-mortems record the classic pattern: the agreement was that a given indicator was for observation only, not for assessment, and then the moment the system came under pressure it was gradually folded into the annual review. The tin opener turned back into a dial. There is a subtler layer of self-scrutiny too, called the Shirky principle: a team whose existence rests on "governing eval defects" depends on those defects continuing to exist, and so it may end up quietly keeping the problem alive. Decoupling is therefore a discipline requiring continuous maintenance: rotate regularly, bring in independent outsiders, and audit what your own eval setup produces. This connects to the institutional layer in G06.
In practice, decoupling is three steps:
- Keep two separate sheets: one is the diagnostic sheet, where the eval tells you where a problem might be; the other is the decision, where a human reads the diagnostic sheet and decides what to change and how. Do not merge them.
- Investigate anomalies, do not settle automatically: when a score suddenly improves or worsens, the first reaction is to ask why and dig in, not to auto-approve or auto-penalize.
- Run a POSIWID health check on a schedule: every so often, ask what your eval has actually been producing lately. Spread out the outputs the last few versions judged excellent and look at them. Did they really get better, or are they just flattering this set of criteria? If the answer is "longer, more obliging, tuned to this particular test", it has stopped being your tool.
One open question: tin openers degrading into dials seems to happen everywhere. Is it reversible? Is there an institutional design that holds an indicator steady in the tin-opener position, for example a rule explicitly forbidding it from being tied to automatic rewards? This is the key gap in turning decoupling from a slogan into an executable institution, and the D7 report lists it as an open problem.
The one-line takeaway: an eval is a health report, not a prize machine; the moment a score carries stakes, the thing being measured starts pleasing the score.
Sources / further reading
- Deming, W. E. Out of the Crisis (1986) (14 points, seven deadly diseases, p.121 Nelson quotation, red bead and funnel experiments); The New Economics (1993) (the "costly myth" wording).
- Ashby, W. R. (1956). An Introduction to Cybernetics, ch.11 "Requisite Variety" (a low-variety regulator cannot suppress a high-variety system).
- Beer, S. (2002). "What is cybernetics?" Kybernetes 31(2):209–219; The Heart of Enterprise (1979) (source of the POSIWID phrase).
- Carter, N. (1991). "Learning to measure performance: the use of indicators in organizations." Public Administration 69(1):85–101 (tin openers vs dials).
- Klein, R. (2025). "The politics of reinvention" (NHS: empirical evidence of tin openers degrading into dials); the Shirky principle — Kelly, K. (2010), The Technium.
- Berenson, R. A. (2016). "If You Can't Measure Performance, Can You Improve It?" JAMA 315(7):645–646.
- Primary sources
research/08(Ashby, Deming, POSIWID, tin opener provenance) andresearch/deep/D7§4 (misquotation forensics and degradation risk).