REACTOR
B10 Laws Laws · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ B10 ]…

Baker Distortion

Cheap, low-noise and controllable, and still a bad metric.

B09 covered crowding out across tasks: grade A and B suffers. There is a sharper question. Even with only one task left, a cheap, precise measure everyone accepts can still be a bad measure. George Baker proved this in 1992 and gave the real criterion for a "good measure." Intuition says a low noise measure is good. Baker says no: good and bad is not about how accurately it measures, it is about whether it points the right way.

Set up the picture first. Every action you take moves two numbers at once: true value (what your boss actually cares about) moves a little, and the performance measure (the number written into the contract) moves a little. An ideal measure is a faithful shadow: whatever action you take, shadow and body move in the same direction and in the same proportion. A bad measure is a funhouse mirror: some actions are stretched in the mirror, others are shrunk. And what you care about is exactly what the mirror shows, because your pay is calculated on that number. So you find yourself picking out exactly the actions that look good in the mirror. Baker turned the distortion of that mirror into a mathematical quantity: the correlation between an action's marginal effect on the measure and its marginal effect on value (written ρ(P′, V′)). Correlation of one and the mirror is flat; the lower it goes, the more warped the funhouse mirror.

FIG.01 What makes a measure good is not low noise, it is alignment EXPLORABLE

The most persuasive thing in the paper is a pair of twin firms (pages 610 to 611). Both pay salespeople commission on revenue, and revenue is equally cheap and equally measurable at both. At Firm A the salesperson's effort only affects how much gets sold, and firm value is also determined only by how much gets sold. So every unit of movement in revenue faithfully corresponds to a movement in firm value: the mirror is flat. Since the mirror is flat, the commission rate can be maxed out and everyone is happy, which economics calls first-best. At Firm B the salesperson's behaviour also affects costs, for instance by promising rush delivery to close a deal, so revenue rises while profit quietly bleeds: revenue's movement no longer faithfully corresponds to value's movement, and the mirror is warped. Once the mirror is warped, the optimal commission rate has to fall below one. Same measurability, opposite conclusions. That controlled-experiment-style example shows that how measurable something is, is not in the criterion at all.

The other classic is the R&D scientist paid per patent (page 609). Patents come easy and hard, and an easy patent is worth as much in the measure as a hard one while being worth far less to the firm. So the scientist over-invests in easy patents and under-invests in hard and important problems. Note that this is not data fabrication. Every patent is real. He is simply picking the actions that look good in the mirror. Baker gives this behaviour a formal definition, gaming: exploiting states where the measure's response exceeds value's response. And the firm's optimal reaction is not moral education, it is lowering the piece rate. The definition of the whole phenomenon on page 600 is equally clean: actions that increase the contractual payment without improving real performance. This is the contract version of Goodhart's law: distortion losses scale continuously with incentive strength. So "do not fully optimize a biased measure" is not advice, it is a property the optimal contract comes with.

Two formal results are enough to remember. First, the optimal commission rate has a closed form (equation 8, page 605): the numerator is one plus the covariance between the measure's marginal response and value's marginal response, the denominator is one plus the variance of the measure's marginal response. The better aligned the direction, the more freely you can give commission; the more crooked, the more you must hold back. Second, page 605 carries a line that often gets skipped: value can fluctuate far more than the measure and none of it matters, as long as those fluctuations do not change what the agent should do. Big noise is not a problem, because noise cannot fool anybody's actions; a crooked direction is fatal, because direction quietly commands actions every day. The conclusion holds under risk neutrality too (the abstract's own words: even if the agent is risk neutral, such contracts generally fail to achieve first-best). So this is not the old story about people fearing risk, it is a pure alignment failure.

The later literature dissects the two defects more finely. Baker's own 2002 sequel formally decomposes measure defects into two orthogonal components, risk (noise) and distortion, each entering the optimal contract in a different way: noise lowers the incentive strength you dare to give, distortion sends the incentive you do give down the wrong path. The accounting scholars Feltham and Xie had already named the two defects and proved their independence back in 1994: a lack of congruity (the measure vector pointing in a different direction from the value vector) distorts the allocation of effort, while noise lowers its intensity. That yields a usable mental formula: how much a measure gets Goodharted is roughly its directional gap times the incentive strength placed on it. It also incidentally explains why weighted composite measures are dangerous: adjusting any one weight rotates the direction of the whole measure vector, and aligning that with a high dimensional true value is close to impossible.

Two more corollaries are worth taking away. First, relative performance evaluation (benchmarking against peers and deducting) is not automatically immune: page 611 proves that it is undistorted only when the benchmark is unaffected by your own actions, otherwise "hold your peers down" becomes a new mirror action. Second, Baker's single-task model and the multitask model (B09) are actually equivalent, and the two papers claim each other in footnotes (Baker page 603 footnote 6, HM page 34 footnote 11): crowding out across tasks and skew within a task are two sides of one coin, and their shared non-trivial conclusion is that the failure has nothing to do with risk.

One open question: the directional gap (congruity) is a clean quantity in theory, but nobody knows how to estimate it from data. Given a benchmark and a real deployment objective, can you measure the directional gap between them before optimization starts, and thereby forecast how far this eval will be Goodharted? The D6 report argues this step is the missing piece that would turn "judging a measure in advance" from a craft into a science, and there is no mature method for it yet.

The one-line takeaway: a bad measure is not one that measures inaccurately, it is one that points the wrong way; and the more stable and trusted a wrong direction is, the more damage it does.

Sources / further reading
  • Baker, G. (1992). "Incentive Contracts and Performance Measurement." JPE 100(3):598–614 (the b* formula eq.8 p.605; definition of gaming p.609; Firms A/B pp.610–611); Baker (2002). JHR 37(4) (orthogonal decomposition of risk and distortion).
  • Feltham, G. A. & Xie, J. (1994). The Accounting Review 69(3):429–453 (naming of congruity and noise).
  • Budde, J. (2007). J. Accounting Research 45(3) (the limits of the balanced scorecard).
  • Gibbons, R. (1998). JEP 12(4) (integrated account with the multitask model).
  • Page-by-page notes in research/03b; deepening in research/deep/D6 §1; reliability-validity mapping in research/06 §4.
Where to next