REACTOR
G08 Defense Defense · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ G08 ]…

Reality as the Ultimate Held-Out

The most cheat-proof exam room is reality itself.

Requires G03 Held-Out & Dynamic Benchmarks Unlocks

The ultimate move is also the plainest: actually put the system out into reality and score it on what it produces in the real world. The reason is simple. The most cheat-proof examination hall is reality itself. Items on a leaderboard can be memorized, ground on and hidden, but whether users actually stayed and whether the job actually got done are outcomes that cannot be faked in advance. The price is that reality's feedback is slow and noisy, and the slowness and noise are themselves the cost.

You learned three anti-cheating moves in G03: hide items, test only on items that appeared after release, and rotate on a schedule. The core of all three is leaving the tested party no way to know in advance what will be tested. But that lesson also poured cold water: hiding items buys time, not immunity. Reality is the most thorough possible time slice. It does not even need to buy time, because a future that has not happened yet physically cannot enter anyone's preparation.

First, why reality is the ultimate held-out set. The key is one property: it cannot be self-reported. Leaderboard results are run and reported by the vendor itself, with the signal source in the hands of the tested party. LIBOR in K04 failed exactly here: a benchmark computed from banks' self-reported rates ended up being collectively suppressed by the self-reporters until it had to be abolished. Real outcomes are the opposite. Whether a user clicked, bought, or came back the next day is an accomplished fact occurring in a world that neither you nor the tested system controls, and nobody can fabricate it from a desk. This is exactly what honest-signalling theory says: whether faking can be prevented splits into two cases. In one, faking simply is not worth it, the cost is too high. In the other, faking is structurally impossible. Real outcomes are the second kind, the hardest class of signal, unfakeable by structure rather than by expense. It also explains a recurring embarrassment. Plenty of research points out that evaluation scores and real usage value are substantially decoupled. Llama 4 is the prototype case: it scored extremely high on a pile of existing benchmarks, while users found on contact that the scores said nothing about how good it was to use. Leaderboards test whether something looks like it can; reality tests whether it actually can.

How to wire reality in: put the system into production and measure what it produces. Run A/B experiments (split users randomly into two groups, one on the new system and one on the old, and compare real behavior), watch outcome metrics in production, look at real user retention. But an old trade-off from B10 is hiding here. Baker divides measures into two kinds. Output measures such as profit or real retention are final outcomes with low distortion, because it is hard to fake them without changing the real thing, but they are noisy, churned by luck, seasonality and the wider environment. Input measures such as click counts or dwell time are process actions with low noise, readable the same day, but with high distortion, far too easy to inflate in ways misaligned with final value. No free lunch: if you want reality's "cannot be self-reported" honesty, you take its noise along with it. If you find it too slow and retreat to instant process proxies, you invite distortion back in. This is also why leaderboards and reality have to be read together. Leaderboards are the low-noise, high-distortion end, with readings available the same day but easily misaligned with real value. Reality is the high-noise, low-distortion end, with readings you have to wait for but which are hard to fake. Neither replaces the other. They occupy opposite ends of an axis, and what you want is a balance in the middle rather than betting everything on either end.

The cost is two words: slow and noisy. Slow means latency. Real retention and long-term consequences, the most valuable signals, inherently require waiting, for users to spend a while with it and for consequences to settle out, while leaderboard scores are instant. If you want reality as your criterion, you have to accept that the verdict arrives long afterwards, and that you must keep making decisions in the meantime. Noisy means noise. Baker's point already said it: the closer an output measure is to the final outcome, the more it is contaminated by luck and the wider environment. Nastier is the optimizer's curse from B13: when you pick the candidate with the highest real-world score from a pile, the score itself has luck mixed into it, not just true ability. You therefore systematically pick the one that luck happened to favour this time, and the true level of the chosen one is, in expectation, below the score it just posted. Noise does not merely make readings wobble, it systematically deceives you at the moment you pick the best. Deming took this all the way long ago, listing management by visible figures alone as one of the deadly diseases and pointing specifically at the fact that the most important quantities are often unknown and unknowable. The reality examination hall is the most honest one there is, and it writes the things that matter most in the slowest, noisiest, hardest-to-read column.

One open question: real-world feedback is slow and noisy, so before the verdict arrives we can only lean on leaderboards and process proxies, the fast and distorted signals. Is there a unified algorithm that fuses reality's slow-but-accurate signals and proxies' fast-but-dirty ones, weighted by their respective latency and noise, so that you neither sit and wait nor get pulled off course by the fast signal? What that really asks is how to trade off optimally between waiting for the truth and keeping to schedule. There is no ready answer today.

The one-line takeaway: reality is the only ultimate examination hall that cannot be faked by self-report; the price is slow and noisy feedback, and having bought honesty you have to learn to read messy handwriting.

Sources / further reading
  • Baker, G. (1992/2002). Distortion vs noise (distortion vs risk). JPE 100(3):598–614; JHR 37(4):728–751 (output measures are low-distortion and high-noise, input measures low-noise and high-distortion; no free lunch).
  • Deming, W.E. (1993). The New Economics (the seven deadly diseases: management by visible figures alone; the most important quantities are often unknown and unknowable).
  • "Evaluating LLM Metrics Through Real-World Capabilities." arXiv:2505.08253 (substantial decoupling between evaluation and real use); the Llama 4 benchmaxxing prototype case.
  • Smith & Winkler (2006). "The Optimizer's Curse." Management Science 52(3):311–322 (picking the highest score systematically overestimates true value).
  • Lachmann, Számadó & Bergstrom (2001). "Cost and conflict in animal signals." PNAS 98:13189 (index / structurally unfakeable signals); Klein (2025), "The politics of reinvention" (tin openers degrading into dials).
  • Primary sources research/deep/D7 §3, §4 (skin in the game, POSIWID, Deming, tin opener) and research/06 §1 (Baker's distortion-noise), §8 (decoupling from real utility).
Where to next