One ruler will always be gamed, so use a set instead. A set of metrics pulling against each other makes "grinding on just one" show up immediately in another. There is a trap here though: adding metrics is not insurance, and a badly built set is easier to fool all at once.
You met surrogation in B07: people sincerely treat "the customer satisfaction survey score" as customer satisfaction itself, without feeling they are cheating. The multitask agent in B09 added a theorem: when someone must split effort between easily measured and hard-to-measure work, rewarding the measurable work heavily pulls their effort out of the unmeasurable work. How a set of metrics holds against both at once is what follows.
The core move is paired metrics: for every metric that can be pushed up on its own, add an opposing metric that gets worse when you push the first. The two pull against each other, so grinding one end shows up at once at the other. Grove's examples in High Output Management are down to earth: measure how many breakfasts were sold, and pair it with how much inventory is left over and how many customer complaints came in; count task volume, and pair it with how many new bugs each version introduced. For a system that genuinely does the job right, both ends look good naturally. For one looking to cut corners, making two numbers that move in opposite directions both look good gets expensive fast.
Why it has to be a set and a single metric will not do has hard evidence from a set of management accounting experiments. Researchers found that tying a bonus to a single metric noticeably worsens the surrogation described above, with people treating the number more and more as the goal itself. Put several metrics side by side and the distortion eases, because several numbers together are a standing reminder that no single number is the whole thing. The most brutal real-world version is Wells Fargo: the cross-sell number started as a proxy for depth of customer relationship, and once it was tied to bonuses and job security, employees first sincerely treated the number as the health of the business. That sincerity did not stay put. It ended in roughly 3.5 million fake accounts, destroying with their own hands the customer relationships the metric was meant to protect. Single metric plus heavy reward can slide all the way from sincere misunderstanding to outright fabrication.
There is a counterintuitive rule for setting weights, and it comes straight from that theorem in B09. When some work is easy to measure and some is hard, the more heavily you reward the easily measured work, the more you are subsidizing withdrawal from the hard-to-measure work. So do not give a metric high weight because it is easy to measure. That is exactly what pushes an agent toward doing only what can be gamed and dropping what genuinely matters but resists measurement. The more asymmetric the measurability, the weaker the incentives should be and the more you should look loosely across several dimensions, rather than betting heavily on a single measurable proxy.
Whether a portfolio resists gaming depends entirely on getting the pairings right. Pairing works because it forces a cheater to fool several dimensions pointing in opposite directions at once. If two metrics can both be raised by the same trick, that is fake balance: one manoeuvre makes both look good. So choose opposites that genuinely pull against each other and whose cheating methods do not transfer. To judge what is wrong with a metric, look at how far the effort it demands departs from real value. That is called distortion. A metric can be cheap, stable and easy to control, and look like a good metric for it. But measured against distortion, it can still be a bad metric, because its distortion is large. That is exactly the criterion in B10. So set weights by alignment with the real goal, not by ease of measurement.
The set also needs a clock. There is a counterintuitive observation: the more metrics there are, the more distortion and gaming an organization tends to show. Multiple dimensions are not insurance in themselves. Following that, metrics have to be managed as moving targets and rotated on a schedule, because any metric dilutes once it has been figured out, in the way described in R01: everyone converges on it and it stops working. This also explains why the balanced scorecard is not a reliable anti-gaming device: it draws a causal chain from leading to lagging indicators, and that chain is usually unvalidated, which can give you false confidence and can let several dimensions share one shortcut and be fooled together by a single move. It is a good tool for communicating strategy, not a safe for resisting gaming. Rotation itself has costs: comparability across years suffers and the tested party has to readjust. Those are the fixed costs a portfolio approach has to accept.
To build a set of fake-resistant metrics, work these four steps:
- Pair: for every metric that can be pushed up alone, add an opposing metric that worsens when you push it (task success rate paired with harmful side-effect rate, resource consumption, hallucination rate).
- Test the pairing: try to think of a cheat that makes both ends look good. If you can think of one, the pair is fake balance. Rebuild it.
- Weight: set weights by alignment with the real goal, not by giving a metric heavy weight because it is easy to measure.
- Rotate: give each metric a service life and swap it out on schedule as a moving target.
One open question: how many metrics should an anti-gaming set optimally have? Too few and coverage has gaps, leaving dead zones where effort can be withdrawn. Too many and, as Meyer records, distortion and gaming rise along with dimensionality, while rotation and cognitive costs run out of control. There is currently no unified framework putting the tension in the pairings, the alignment of the weights and the rhythm of rotation into a single calculation of how much gaming I will tolerate and how much comparability and complexity I will pay for. This is also one of the dials the C03 design lab exists to turn.
The one-line takeaway: a good set of metrics makes "fooling every dimension at once" more expensive than doing the job right; a bad set is just several copies of the same hole.
Sources / further reading
- Choi, J., Hecht, G. W. & Tayler, W. B. (2012). "Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation." The Accounting Review 87(4):1135–1163; (2013). Journal of Accounting Research 51(1):105–133 (single metrics worsen and multiple metrics mitigate surrogation).
- Harris, M. & Tayler, W. B. (2019). "Don't Let Metrics Undermine Your Business." HBR (the Wells Fargo case).
- Holmström, B. & Milgrom, P. (1991). "Multitask Principal-Agent Analyses." JLEO 7:24–52 (measurability asymmetry, the equal-compensation principle).
- Baker, G. P. (1992). "Incentive Contracts and Performance Measurement." JPE 100(3):598–614 (decomposition into distortion and noise).
- Meyer, M. W. (2003). Rethinking Performance Measurement: Beyond the Balanced Scorecard (more metrics, more gaming; moving targets).
- Nørreklit, H. (2000). Management Accounting Research 11(1):65–88; Smith, M. J. (2002). JMAR 14:119; Grove, A. (1983). High Output Management (paired metrics).
- Eisenstein et al. (2023). "Helping or Herding? Reward Model Ensembles Mitigate but Do Not Eliminate Reward Hacking." arXiv:2312.09244 (same-model ensembles get fooled together).
- Primary sources
research/06§1,research/03b,research/03c(surrogation) andresearch/deep/D7§4 (the scorecard critique, Meyer).