A story management courses have told for seventy years: in the 1920s, at the Hawthorne works outside Chicago, researchers turned up the lights on the shop floor and output rose; they turned them down and output rose again; they changed nothing at all and it still rose. Conclusion: what mattered was not the lighting but the fact of being paid attention to. That is the Hawthorne effect, probably the most famous evidence that being observed changes behaviour. It also has a second half far fewer people know: once the original data was dug out and checked, the story collapsed. This is not trivia. The whole red branch is about measurement changing the measured, so the most famous showcase for that claim has to survive the same scrutiny.
From 1924 to 1932 the Western Electric Hawthorne works in Cicero, Illinois ran a series of studies: the illumination experiments (from 1924, in collaboration with the National Research Council) examined the relation between light levels and output; the relay assembly test room (1927 to 1932) put five women workers in a separate room and varied hours, breaks and pay schemes one at a time. The authoritative report is Roethlisberger and Dickson's Management and the Worker (1939), on which Harvard's Elton Mayo built his theory, and the entire management doctrine of the "human relations school" rests on this story. One piece of trivia along the way: the name "Hawthorne effect" was not coined by the original researchers, it was minted by French in 1953 and popularised by Landsberger's 1958 reanalysis, so pinning it retroactively on Mayo is an anachronism. What was overturned is only this case, and reactivity itself is untouched. Tracing it shows two things: classics also need checking, and the real scale of reactivity is institutional, not psychological.
There were three waves of reanalysis, each more damaging than the last. First, Franke and Kaul in 1978 ran the first statistical reassessment and found that output changes were better explained by managerial discipline (replacing low-output workers), the economic pressure of the Depression, and rest and pay arrangements, rather than by being observed. Second, Jones in 1992 reanalysed the relay assembly room data with modern econometrics and found no robust evidence. The third wave was the decisive blow: the raw illumination data had been thought destroyed and never formally analysed, and Levitt and List recovered the archive from two libraries in 2011, concluding on reanalysis that the striking patterns described in the literature had no support in the records, their phrase being "entirely fictional." Output fluctuations in fact tracked two utterly ordinary rhythms: day of the week (Monday and Friday) and the pay cycle. Control for those two and the difference between treatment and control groups essentially vanishes, leaving only the faintest trace. Even the most dramatic detail did not survive: "workers still raising output at near-moonlight illumination levels" also has no clean evidence in the archive; and the early gains in the relay room came at least partly from two low-output workers being replaced, an explanation with far more substance than the lighting. Earlier still, Adair pointed out in 1984 that the effect is defined inconsistently and operationalised in all sorts of ways across the literature, and that reckoning has been cited around 1900 times.
Evidence-based medicine delivered a procedural verdict. The way that field works is the systematic review: gather all the studies on a question, grade them by evidence quality, then conclude. McCambridge, Witton and Elbourne reviewed 19 studies in 2014 and concluded that the "Hawthorne effect" is too vague a concept with unclear mechanisms, recommending the more precise "research participation effects" instead; that paper, cited around 3,800 times, is in substance a proposal to retire the name. Subsequent tests point one way: McCarney and colleagues' rare randomised controlled trial in 2007 found limited evidence; a 2019 randomised trial designed specifically to induce a Hawthorne effect measured nothing at all; a 2022 meta-analysis in primary care leaned toward keeping the folk name but renamed the mechanism "participant reactivity." One telling detail: in 2017 some scholars proposed that medical education research simply switch to the word reactivity, converging with the red branch's own vocabulary. The trend is clear: the stricter the design, the harder a strong effect is to reproduce.
"Being observed changes behaviour" is really a family of effects with different mechanisms and uneven evidence, and lumping them all under "Hawthorne" mixes good evidence with bad. The John Henry effect (named in 1972, after the folk-ballad steel driver who raced a steam drill and worked himself to death) is about the control group: they know they are being compared, so they try extra hard. Demand characteristics (Orne 1962) is about participants working out what the experimenter wants and then performing accordingly. There are also novelty effects, the Pygmalion effect of teacher expectations, and the placebo effect. Citing any one of these should mean checking its own evidence separately, not borrowing blanket endorsement from the umbrella term "Hawthorne."
This case explains why the red branch has the shape it does. Precisely because "the psychological shiver of being observed" is so hard to reproduce robustly, Espeland and Sauder recast reactivity from a laboratory contaminant into an observable macro process at the institutional level: what they studied was not the heart rates of women workers but how rankings reorganised the entire field of legal education (that lesson is R01). Placing the Hawthorne myth next to law school rankings highlights the difference between two kinds of reactivity: on one side weak, hard-to-replicate individual responses, on the other robust, sustained institutional responses. What drives the latter is not the feeling of being watched but the structural incentives imposed by a commensurable, public, zero-sum measure.
One open question: if "observation changes behaviour" has thin evidence at the individual psychological level but is extremely robust at the institutional level (rankings, audits, indicators, the content of the rest of this branch), what determines the dividing line? The D1 report's candidate answer is incentive structure: reactivity only gets sustained fuel when the reading is wired to money, reputation and survival.
The one-line takeaway: a glance may not change you, but when the eye watching you is wired to your livelihood, change is certain, and it happens in institutions, not in heartbeats.
Sources / further reading
- Roethlisberger, F. J. & Dickson, W. J. (1939). Management and the Worker. Harvard UP; Mayo, E. (1933) (the original report and its theorisation).
- Levitt, S. D. & List, J. A. (2011). "Was There Really a Hawthorne Effect at the Hawthorne Plant?" AEJ: Applied Economics 3(1):224–238.
- Franke & Kaul (1978). ASR 43(5); Jones (1992). AJS 98(3):451–468; Adair (1984). J. Applied Psychology.
- McCambridge, J., Witton, J. & Elbourne, D. R. (2014). "Systematic Review of the Hawthorne Effect." J. Clinical Epidemiology 67(3):267–277 (the verdict to retire the name); McCarney et al. (2007); Berkhout et al. (2022); Paradis & Sutkin (2017).
- Chiesa & Hobbs (2008). EJSP 38(1); Saretsky (1972) (the John Henry effect); Orne (1962) (demand characteristics).
- The comparison family of effects and the latest re-checks:
research/deep/D1§7; falsification details:research/07§3.