REACTOR
B09 Laws Laws · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ B09 ]…

Multitask Agency

Why teachers get a flat salary: there is a theorem for it.

Kerr named the common disease of organizations: hoping for good teaching while rewarding only publication (B06). The natural reaction is to hang the target properly and start grading teaching too, and surely the problem is solved. That intuition is wrong. In 1991 the economists Bengt Holmström and Paul Milgrom proved mathematically that as long as the important things cannot be measured accurately, the harder you grade the measurable things, the faster the unmeasurable part gets drained. Push that logic to its extreme and there is only one conclusion: in some cases the best contract is to pay no performance money at all and give a fixed salary. Teachers being on fixed salaries is not a leftover of lazy administration, it is the precise optimum. The theorem later became one of the core contributions for which Holmström won the Nobel Prize in economics in 2016.

Start with an everyday version. You hire decorators to paint your walls and pay by the square metre. Square metres are easy to count, so the walls get painted fast; edges, corners, how many coats of primer, the details that decide how long the paint lasts cannot be priced by the metre, so they all get skipped. You did not hire lazy people. You hired smart people who respond sensitively to a price list, and they moved all of a day's limited energy onto the dimension that converts into money. The multitask model is the physics of that transfer machine.

FIG.01 The multitask model: a theorem about the fixed salary EXPLORABLE

Say the skeleton of the model out loud and there is really only one load-bearing assumption: a person's energy is a single total. Do more of this and you do less of that, which economics calls effort being substitutable across tasks. Accept that one, and the conclusions run downhill from there.

A bonus has two functions in this world, and most people see only the first. Function one is the accelerator: pay more and people work harder overall. Function two is the steering wheel: whichever task the pay hangs on is where the energy flows. Page 25 of the original spells it out, that incentive pay does not only allocate risk and motivate effort, it also directs how an agent divides attention among their various duties. The trouble is that the steering wheel only recognizes what is measurable. If a task cannot be measured, its marginal reward becomes zero. Once the marginal reward hits zero, energy naturally drains away from it.

The teacher case (pages 32 to 33) pushes the mechanism to its limit. Teaching basic skills can be picked up by a test; teaching independent thinking cannot. If you want teachers to spend more energy on the part that cannot be measured, only one route remains: lower the reward on the part that can, and reduce the opportunity cost of the attention being pulled away. Follow that logic to the end and you get Proposition 1 on page 34: when the important task is entirely unmeasurable, the optimal contract is a fixed wage with zero incentive, even if the teacher has no fear of income variation at all. That "even if risk neutral" is the cutting edge of the whole paper. Economics used to explain weak incentives by "people are risk averse, so use less commission." This theorem says that even if nobody feared risk, a steering wheel pointed the wrong way is worse than no steering wheel.

So the core result is worth memorizing (page 26): whether a task should get strong incentives does not depend on how measurable that task is, it depends on how measurable its neighbours are. You never grade one task. You grade one person's whole timetable.

The theorem carries a string of organizational corollaries, each of which can be matched to reality. Weak incentives in the civil service are institutional rationality, not bureaucratic sloth: the dimensions that matter most in public service (fairness, integrity, long term consequences) are exactly the hardest to measure. Piece rates fail especially easily in large hierarchies (the original wording, page 34), because the bigger the organization, the deeper the measurement gap between quantity and quality. Conversely, the most constructive corollary is to treat job design itself as an incentive instrument: bundle tasks by measurability, cluster the measurable ones into one role with strong incentives and the hard-to-measure ones into another role on flat pay, and do not make the same person carry both kinds. Dewatripont, Jewitt and Tirole formalized that prescription in 2000 under the name task clustering.

The paper also buries two lines that are easy to miss. One is asset ownership (pages 26 to 27): a franchisee bearing their own profit and loss gets the strongest incentives, a company-store manager gets weak incentives plus corporate procedure, and both arrangements are internally consistent. The reason is that ownership itself determines which returns a contract can perceive; the more of those returns can be perceived, the stronger the incentive should be. The other is the claim on page 27: constraints are a substitute for performance incentives. When you cannot use a bonus to steer direction, prohibitive rules (banning outside work, say) can take over part of the incentive function, which explains why heavy regulation and weak incentives always come in pairs in mature organizations.

That line has run on since. Gibbons's 1998 survey merged it with Baker's distortion model from the next lesson (B10) into a single sentence: what you measure is all you get. Holmström's 2017 Nobel lecture tied it off himself: the core of incentive design is not to measure and reward more, it is to align the whole incentive package with the true objective. Bénabou and Tirole pushed one step further to the market level in 2016: when an industry competes for talent with high pay, over-investment in the measurable dimensions is further amplified by competition. Amplified across a whole industry, "quantity over quality" becomes that industry's equilibrium. Under that equilibrium, the KPI arms race in finance and tech is not individual deviance, it is an equilibrium outcome. A batch of work since 2024 also suggests that when the people being graded can see each other's measures, peer comparison enlarges the crowding-out effect further.

One open question: the load-bearing assumption of the theorem is that effort is substitutable, since a person's day has only twenty-four hours, and for people that assumption clearly holds. For a model with a fixed parameter count, do training compute and representational capacity crowd each other out across capability dimensions in the same way? The crowding evidence in Y02 says the direction is right, but how to define "a model's total effort" and how large the elasticity of substitution is still have no Holmström-style theorem.

The one-line takeaway: a bonus is not only an accelerator but a steering wheel; when the important thing cannot be measured accurately, weak or even zero incentives are the optimal design.

Sources / further reading
  • Holmström, B. & Milgrom, P. (1991). "Multitask Principal-Agent Analyses." JLEO 7:24–52 (Proposition 1 p.34; core quotations pp.25–26, 32–33).
  • Milgrom, P. & Roberts, J. (1992). Economics, Organization and Management (naming of the equal compensation principle).
  • Dewatripont, Jewitt & Tirole (2000). European Economic Review 44 (task clustering).
  • Gibbons, R. (1998). JEP 12(4) (survey); Holmström, B. (2017). AER 107(7) (Nobel lecture summation); Bénabou & Tirole (2016). JPE 124(2) (bonus culture).
  • Page-by-page notes in research/03b; deepening in research/deep/D6 §1; briefing research/12.
Where to next