In the Hanoi rat hunt of B04, the government wanted fewer rats in the city, but it was not going into the sewers itself, so it hired trappers; it could not see what they did down there and could only count the tails handed in. That triangle (I want A, I hire you to do it, but all I can see is B) has a formal name: the principal-agent problem. The principal is the party paying for the outcome, the agent is the party taking the money and doing the work, their goals are not fully aligned and their information is asymmetric. Every story in the blue branch has this as its skeleton. A measure is the searchlight the principal uses to pierce the information asymmetry: without it, the principal cannot see what the agent is doing in the dark. But the moment the searchlight comes on, the agent sees where the beam lands too, and starts adjusting its moves to follow the light. Wherever the beam falls gets gamed.
There are two kinds of information asymmetry, striking at different times, and the memory anchors are "before" and "after." The first is moral hazard, arising after the contract is signed: you cannot see how hard the other party works, only an outcome mixed with luck. Whether the tests a doctor orders are necessary, whether a contractor uses solid materials, whether a model really learned anything during training, all belong here, and the technical term is hidden action. The second is adverse selection, arising before the contract is signed: you do not know what kind of person you are dealing with. An employer cannot tell the diligent applicant from the idle one, an insurer cannot tell the healthy customer from the one who is already ill, and the technical term is hidden type.
Akerlof took the destructive power of adverse selection to its limit in 1970 with the used car market, and the story is worth remembering in full. The seller knows whether the car is good or bad, the buyer does not, so the buyer will only pay for "average market quality." That average price is a loss for owners of good cars, so good cars gradually leave the market; once they leave, average quality falls, so buyers bid lower again, driving out the next tier of decent cars. A few rounds of this and only bad cars remain (lemons in American slang), or the market disappears entirely. What is frightening about the paper is the scale of the conclusion: information asymmetry does not just make transactions expensive, it can kill the market itself. Transposed to evaluation: when buyers (users, investors) cannot distinguish real model quality and can only price at "average expectation," the vendors doing serious safety work are subsidizing the vendors who talk big, and the engine of bad money driving out good is already running.
Faced with these two asymmetries, the principal's instinct is to measure more and attach more pay, and this is exactly where theory hits the brake. Holmström's 1979 informativeness principle says that whether a signal is worth writing into a contract depends only on whether it provides incremental information about what the agent did. That principle is most often read backwards. The popular version, "anything measurable should be graded," is its opposite: a signal that carries no information about the unmeasurable dimensions while being highly sensitive to the measurable ones will, once in the contract, actively drain effort toward the measurable (the mechanism of B09). The criterion was never measurability, it is information increment: a new measure highly correlated with an existing score has an information increment of roughly zero, and putting it in the contract only adds surface for gaming.
Fifty years of literature have accumulated two genuine institutional exits, both written in step four of the module, and each deserves a sentence. Exit one is subjective measurement. When every objective measure gets gamed, letting a supervisor score on overall impression can recover part of what is lost. But that impression score is, in the end, a subjective judgment: there is no actionable objective standard, so a court cannot enforce it as a contract. The only thing that makes it work is long term reputation, self-enforcing: if you score arbitrarily today, nobody trusts your assessments tomorrow. The technical term is a relational contract (Baker, Gibbons and Murphy 1994). Exit two is strategic opacity. The agent watches the weight table and piles effort onto whichever items carry the highest weight; publishing the weights just tells the agent which dimensions it can safely neglect. So the principal's best response is often not to publish the weights: the agent cannot guess which dimension weighs most, and the safest play is to spread effort evenly across all of them (Ederer, Holden and Meyer 2018). Secrecy is not administrative laziness, it is mechanism design derivable from the first order conditions of an optimization. The prescription British public management scholars wrote for health targets, random audits with the assessment definitions published only after the fact (K02 tells that history), acquires rigorous microfoundations here.
One open question: both institutional exits presuppose that the agent is a person. When "supervisory judgment" is replaced by an LLM judge, are the mechanisms subjective measurement relies on to self-enforce (reputation, repeated play, the long term cost of being caught lying) still there? When the party being graded is a model that can read every assessment rule ever published, how long can strategic opacity hold? Nobody has formally joined up this interface between principal-agent theory and AI alignment.
The one-line takeaway: a measure is a searchlight into the dark places of information, and the people in the dark always see the edge of the beam before you do.
Sources / further reading
- Ross, S. A. (1973). AER 63(2); Mitnick, B. (1973) (independent origins); Jensen, M. C. & Meckling, W. H. (1976). JFE 3(4):305–360 (the p.308 quotation).
- Akerlof, G. A. (1970). "The Market for 'Lemons'." QJE 84(3) (adverse selection); Holmström, B. (1979). "Moral Hazard and Observability." Bell J. of Economics (the informativeness principle).
- Feltham, G. A. & Xie, J. (1994). The Accounting Review 69(3):429–453 (the congruity and noise decomposition).
- Baker, Gibbons & Murphy (1994). QJE 109(4) (subjective measures and relational contracts); Ederer, Holden & Meyer (2018). RAND (strategic opacity); Gibbons, R. (1998). JEP 12(4) (survey).
- Briefing in
research/12; deepening inresearch/deep/D6§1 andresearch/deep/D2§6.