Atlas
All 69 nodes, the core glossary, and a quick-reference table of corrections. This course did the legwork of tracing every popular claim back to what was actually said.
Reactivity
Rankings don't just reflect reality. They help remake it.
R02Self-Fulfilling Prophecy
How a false definition grows into a true fact.
R03Commensuration
All of American legal education, squeezed onto three pages.
R04Discipline & Surveillance
Why organizations cannot buffer themselves from a ranking.
R05Legibility
To see you clearly, the state first makes you simpler.
R06Looping Effects
Classifying people changes the people being classified.
R07Performativity
A theory is not a camera. It is an engine.
R08Reflexivity (Soros)
Prices reflect the fundamentals, and also shape them.
R09Audit Society
Making the organization auditable becomes the point in itself.
R10Matthew & MusicLab
The best song still lives or dies on luck.
R11Hawthorne & Its Debunking
Reactivity's flagship case does not survive its own raw data.
R12Mechanical Objectivity
Objectivity was not discovered. Institutions needed it into being.
R13Boycott, Exit & Plurality
Two years after the mass boycott, the ranking was untouched.
R14The Emotional Machine of Rankings
Anxiety, seduction, resistance: how a ranking moves into your head.
R15The Birth of Statistical Thinking
The normal person is a nineteenth-century invention.
Goodhart's Law
How one small footnote came to rule the world.
B02Campbell's Law
The man who warned about metrics was measurement's own champion.
B03Lucas Critique
The moment you act on a regularity, people change their behavior.
B04Cobra Effect
Pay a bounty for cobras, and people start breeding cobras.
B05McNamara Fallacy
Whatever cannot be measured ends up declared nonexistent.
B06Rewarding A, Hoping for B
Hoping for good teaching, rewarding only publication.
B07Surrogation
Why the people doing it never feel like they are cheating.
B08Goodhart Taxonomy
One slogan, four completely different mechanisms.
B09Multitask Agency
Why teachers get a flat salary: there is a theorem for it.
B10Baker Distortion
Cheap, low-noise and controllable, and still a bad metric.
B11Requisite Variety
Only variety can absorb variety.
B12Principal-Agent
You want a good agent. The optimizer only wants the eval to pass.
B13Optimizer's Curse
The option you picked is worth less than its own estimate said.
B14When Metrics Work
When is teaching to the test just teaching? The law has edges.
B15Ridgway 1956: The Forgotten Prequel
Someone wrote all of this down in 1956, then it was forgotten for sixty years.
B16Signals & Cheap Talk
Talk is free, which is exactly why promises are cheap.
B17Regression to the Mean
The leader falls back with no conspiracy involved. The math does it alone.
Overfitting
Your training set is your eval.
Y02Adaptive Overfitting
Look at the same test set often enough and you are training on it.
Y03Contamination & Saturation
The questions leaked into training, and the questions may be bad anyway.
Y04Reward Hacking
The boat-racing AI stopped racing and spun in circles for points.
Y05Nearest Unblocked Strategy
Why every patch to an eval just gets routed around.
Y06Mesa-Optimization
What you trained may turn out to be another optimizer.
Y07RM Overoptimization
The cleanest experimental curve Goodhart's law has ever had.
Y08Sycophancy
Reward human approval, get flattery.
Y09Performative Prediction
Predictions change the distribution: reactivity, written as math.
Y10Strategic Classification
Publish the scoring rule and people reshape themselves to fit it.
Y11Leaderboard Illusion
Launch day, straight to the top. How much of that score is test prep?
Y12Evaluation Awareness
A model that knows it is being watched is not the real model.
Y13Contamination Forensics
You cannot even prove that a score is clean.
Y14Agent Eval Frontier
What gets tested is no longer an answer, but a chain of actions.
Y15The Science of Evals
Running evals on the evals.
Y16RLHF in Plain Words
Before you train an AI to please people, you have to turn liking into a number.
Y17How Benchmarks Get Made
Behind every exam paper there is someone working against a deadline.
Honest Signals & Trade-offs
A signal is honest when it is hard to fake. But is costly enough?
G02Decoupling
Never let the eval score drive your iteration directly.
G03Held-Out & Dynamic Benchmarks
Keep the optimization loop permanently behind your question bank.
G04Metric Portfolios & Rotation
One ruler always gets gamed. What about a whole set of them?
G05Anti-Gaming Mechanism Design
Make cheating cost more than actually being good.
G06Institutional Defenses
The ruler for rulers: who audits the auditors?
G07Red-Teaming & Adversarial Testing
Hire a crew of professional troublemakers to break your eval.
G08Reality as the Ultimate Held-Out
The most cheat-proof exam room is reality itself.
Education: Teaching to the Test
The scores went up. The learning did not.
K02Governance: Targets & Terror
The four-hour miracle on the waiting list.
K03Academia: Citation Arms Race
The impact factor is both the biggest victim and the accomplice.
K04Markets: Ratings & Fake Reviews
The assembly line behind five-star reviews.
K05Medicine: Report Cards & Cherry-Picking
Publish surgical report cards and surgeons start picking patients.
K06Planned Economy: Death by Quota
Judge a nail factory by weight and you get one giant nail.
Concept Map
Five names that are not five names for one thing.
C02The Seven Levers
Predict how badly any ranking or eval will distort.
C03Anti-Goodhart Lab
If every metric gets gamed, what design survives it?
C07Your Eval Postmortem
Finding that your eval changed nothing is worth more than most evals.
C08A Ranking Reader Survival Guide
Next time you meet a ranking, ask these seven questions first.
- Reactivity 反应性
- 人因被测量、观察、评估而改变行为,本课的地基(Espeland & Sauder)。
- Commensuration 可通约化
- 把异质的质压成共享一把标尺的量;主动创造可比性,而非发现相似。
- Self-fulfilling prophecy 自我实现预言
- 起初为假的定义,经由改变行为把自己变真(Merton)。
- Performativity 表演性
- 理论不是照相机而是引擎;使用使世界向理论收敛(Barnesian)或背离(counter)。
- Looping effects 循环效应
- 分类改变被分类者,被分类者反过来改变类别;人的种类是移动的靶(Hacking)。
- Legibility 可读性
- 为治理而把社会简化成可测形式;操纵的前提(Scott)。
- Goodhart 定律
- 任何统计规律一旦用于调控就崩塌;流行的'measure/target'版实为 Strathern。
- Campbell 定律
- 指标越用于高利害决策,越腐蚀它本要监测的过程。
- Requisite variety 必要多样性
- 只有多样性能消灭多样性;单一指标约束不住复杂系统(Ashby)。
- Surrogation 代理指标替代
- 把衡量目标的指标当成目标本身的认知替换,Goodhart 的心理地基。
- Reward hacking 奖励破解
- 优化器把奖励拉满却不做你想要的事;Goodhart 在梯度下降里的重演。
- Nearest unblocked strategy 最近未堵策略
- 打一个补丁,优化器就找到最近的绕行捷径;这是修 eval 总显得机械的原因。
- Adaptive overfitting 适应性过拟合
- 反复用同一测试集做选择,就把它的特异性学进模型:看也是训练。
- Performative prediction 表演性预测
- 预测改变数据分布;重训练收敛到'自造世界'而非最优(Perdomo)。
- Optimizer's curse 优化者诅咒
- 从多候选择优即高估;选中者的真值期望必低于其估计。
- Trade-offs 权衡(诚实信号现代版)
- 诚实由质量依赖的净收益差维持;成本差只是其中一种实现,均衡成本可为零(Számadó et al. 2023/2026)。
- Strategic classification 策略性分类
- 公布分类器,被分类者照规则改造自己;想诱导真改进必须解因果推断问题(Miller-Milli-Hardt 2020)。
- Evaluation awareness 评测觉知
- 模型识别出自己在被测并改变行为;2025 年已有因果证据,读数本身被污染。
- Proxy failure 代理失效
- 跨神经科学/经济学/生态学的统一命名,把成瘾、孔雀尾与 Goodhart 收进同一机制(John et al. 2024, BBS)。
- Good regulator theorem 好调节器定理
- 好的调节器必是被调节系统的(同态)模型;scalar eval = 没有模型的调节(Conant & Ashby 1970)。
- Mechanical objectivity 机械客观性
- 遵循公开规则、排除个人判断的客观性:谁来算结果都一样;不信任裁量的社会转而信任程序与数字(Porter)。
- L'homme moyen 平均人
- Quetelet 把误差曲线搬到人身上:正中央从杂音翻转成理想与标准,「正常/不正常」这条线由此被画出。
- Cheap talk 空口白话
- 零成本的话能传多少真,由双方利益一致程度单调决定;维持诚实所需成本正比于利益冲突(Crawford & Sobel)。
- Regression to the mean 向均值回归
- 极端成绩里掺着运气,再测一次运气散去分数自然回落;先扣回归,再谈反应性(Galton)。
- RLHF 从人类反馈做强化学习
- 人类二选一→奖励模型→拿分当靶子拧策略;三步连续降维,模型最后追的是刻度不是你心里的「好」。
- Bradley-Terry 配对比较模型
- 假设每个选手有隐含实力值,从大量两两胜负里反推;竞技场排名与 RLHF 奖励模型的共同统计骨架。
- Benchmark saturation 基准饱和
- 顶尖模型间失去可统计区分的分辨力;随基准年龄温和上升,公开与私有题库无显著差别。
- Red-teaming 红队
- 雇捣蛋鬼抢在优化器前头找最近的未堵出口;只能证明「这里能破」,永远证不了「哪里都破不了」。
- Temporal split 时间切分
- 只考模型训练截止日之后出现的题,没见过未来就没法背;把防作弊从抬成本换成讲因果。
- POSIWID
- 系统的目的就是它实际所做的,不是它宣称的意图;短语确为 Beer 所述,缩写为后人定型。
- Impact factor 影响因子
- 某刊前两年论文的当年篇均被引;分子分母都留有套利面,DORA 主张不得以它评单篇与个人。
- Selection effect 选择效应(成绩单效应)
- 不改自身水平,改被测总体:医生回避高危病人,账面全真,比造假更难抓(Dranove 2003)。
- Success indicator problem 成功指标问题
- 单一总量指标下的方向性扭曲与品种坍缩:按吨造厚、按面积造薄,指标没写的等于死刑(Nove 1958)。
- 裁判的利益相关度(第八杠杆)
- 打分机构的收入、估值或数据来源是否依赖被打分者;与前七根杠杆正交,裁判可中立也可不中立。
A hidden theme of this course: the popular version is usually not the original. Every line below is documented in full inside the lessons.
| Popular claim | Correction | See |
|---|---|---|
| “When a measure becomes a target…” | 不是 Goodhart 说的。是 Strathern (1997, p.308) 措辞、Hoskin (1996) 命名。 | B01 / R09 |
| 德里眼镜蛇故事 | 无一手史料,疑为都市传说。有档案的是河内老鼠悬赏 (1902, Michael Vann)。 | B04 |
| McNamara 四步引文 | 实为 Yankelovich 1971 演讲;Handy 1994 误归给 McNamara。 | B05 |
| 霍桑效应(强版本) | 原始照明数据经 Levitt & List (2011) 重分析判为 'entirely fictional'。 | R11 |
| Ashby “absorb variety” | 原文是 'only variety can destroy variety';absorb 是 Beer 的改写。 | B11 |
| “不能度量就不能管理”≈Deming | 反了。Deming 视其为要破除的谬误;'最重要数字不可知'是他转引 Nelson。 | G02 |
| “Grafen 证明了障碍原则” | 该流行叙事被 Penn & Számadó (2020) 论证为对模型的误读。 | G01 |
| NUS 归于 Alex Turner | 更准的溯源是 Yudkowsky / Arbital(约 2015)。 | Y05 |
| Pygmalion 效应是铁证 | 效应主要限一二年级、复制不稳定,被 Thorndike 1968 批评。 | R02 |
| 苏联钉子厂 | 寓言级证据,无一手史料;巨钉漫画的《鳄鱼》刊期从未被可靠定位。机制真实但别当史实引用。 | K06 |
| 家族起点是 Campbell / Goodhart | 需前推:Ridgway 1956 (ASQ) 是最早的成文系统综述,早 19 年;三方原文互不引用,大概率独立发现。 | B01 / B15 |
| Goodhart 与 Lucas 独立发现 | 需收紧:Chrystal & Mizen (2003) 裁定「若两者等价,Lucas 几乎肯定先说」。 | B01 / B03 |
| Campbell 1976 与 1979 是两说 | 同一文本:油印本 (1976) = 期刊版 (1979),措辞一致。 | B02 |
| E&S 2007 提出「四机制」 | 2007 原文只有 2 机制 + 3 效果;narrative 出自 Espeland 2015,reverse engineering 与 emotional attachments 出自 Espeland 2016 (HSR)。 | R01 / R14 |
| MusicLab 证明成功纯随机 | 误读:2008 反转实验恰证质量设定边界,最好的歌能从人为打压中恢复。 | R10 |
| 霍桑效应作为统一效应成立 | 连概念本身都被判不成立;McCambridge (2014) 主张弃名,改用 research participation effects。 | R11 |
| 污染 = 分数线性虚高 | 路径依赖:同一次泄漏既可致命,也可被后续训练洗掉(Schaeffer 2026 的污染悖论)。 | Y03 |
| 反复刷同一榜必然严重过拟合 | 经验裁决温和得多:Kaggle 百场「little evidence of substantial overfitting」,机制是模型相似性。 | Y02 |
| winner's curse 出自 Thaler | 概念源自 Capen, Clapp & Campbell (1971) 三位石油工程师;Thaler 1988 是通俗化者。 | B13 |
| 好调节器定理证明需要世界模型 | 「model」只是同态映射;且因果方向常被讲反,原证明有技术空隙(Scholten 2010)。 | B11 |
| POSIWID 是后人讹传 | 这次不是讹传:短语确为 Beer 所述,仅缩写为后人定型。 | G02 |
| cost differential 是诚实信号的终点 | 已被 2023/2026 的 trade-offs 超越:成本可为零,关键是质量依赖的净收益差。 | G01 |
| 昂贵信号是诚实的必要条件 | 错:昂贵只在利益冲突时才需要,且需要量正比于冲突(cheap talk 传统)。 | G01 / B16 |
| 操演性 = 经济学总能自证 | 抹掉了 MacKenzie 四分类学,尤其方向相反的 counterperformativity。 | R07 |
| 「榜首回落」证明反应性 | 向均值回归足以单独解释;须先扣除回归再归因 Goodhart。 | B13 / B17 |
| 统计类别是建构的,所以是假的 | Hacking 明确反对:互动类既是建构的又是真实的,「建构=虚构」是对循环效应最常见的误用。 | R15 |
| Ridgway 拼作 Ridgeway | 误拼多了一个 e,连一些正式出版物都拼错;正确拼法是 Ridgway。 | B15 |
| 家族警句只有那句 measure/target | 最早最形象的警句是 Ridgway 1956 的青霉素比喻「治病的药有时比病更糟」,却几乎从未被通俗文章引用。 | B15 |
| Secrist:商业世界正走向平庸的胜利 | 1933 年整本书把向均值回归误当真实经济规律,成了统计学著名反面教材。 | B17 |
| 《体育画报》封面诅咒 | 上封面正因刚打出极端高峰,之后回落是纯回归;诅咒不存在,存在的是给回归编因果故事。 | B17 |
| RLHF 教模型对齐人类价值观 | 它优化的是「人类当场更喜欢哪个回答」;偏好模型常奖励迎合而非真话,谄媚被放大(GPT-4o 2025-04 回滚)。 | Y16 |
| 逼近满分=模型快把任务做完了 | MMLU 约 6.49% 错题把上限钉在 93.5% 附近:撞的是标注天花板与糙设计,不是能力天花板。 | Y17 |
| 把题藏起来,分数就是干净的 | 公开与私有测试集饱和率无统计显著差异;污染检测方法本身站不住,阴性不等于干净。 | G03 |
| 指标越多越抗博弈 | Meyer 2003 观察相反:指标越多扭曲往往越多;制衡只在各维正交、作弊手段不相通时成立。 | G04 |
| 设计得够好就能造出不可博弈的指标 | Gibbard–Satterthwaite 与 Skalse 等证明一般情形不可能;只能在四条退路上抬高操纵成本。 | G05 |
| 旧榜有病,换套更透明的新榜就好 | Doing Business 换成 B-READY 照被质疑「过时且不公」;US News 改法后博弈照旧。后果还在,改方法只是换战场。 | G06 |
| 过了红队测试就是安全的 | 红队只给下界:证明「这里能破」,证不了「哪里都破不了」;sandbagging 与评测觉知还在掏空地基。 | G07 |
| 上线跑 A/B 看真实数据就客观了 | 现实防的是自报造假,防不了选错代理;现实指标一旦成为目标照样滑回 Goodhart。 | G08 |
| Campbell 举过中国科举的例子 | 1979 原文全文无 China、无 imperial examination,科举例子是后人附会;「Campbell's Law」之名亦系后人所加。 | K01 |
| Bevan & Hood 证明了目标制失败 | 过度简化:受考核维度的改善是真实的,论点是改善与博弈并存、审计太弱无法区分比例。 | K02 |
| 豆瓣「锁分」/「评分皆水军」 | 法院认定锁分证据不足;水军攻击另有分布铁证。攻防双向,单方受害叙事是选择性引用。 | K04 |
| 成绩单证明公开透明是错的 | 把双向证据读成单向:回避高危是实锤,真实质量改进同样成立(Kolstad 2013);净福利取决于风险调整。 | K05 |
| 榜单出问题,改改评分方法就能修好 | 改良指标不等于消除反应性:修订本身立刻成为被博弈对象,指标与被测者共同演化。 | C08 |