REACTOR
R10 Sociology Sociology · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ R10 ]…

Matthew & MusicLab

The best song still lives or dies on luck.

"For unto every one that hath shall be given, and he shall have abundance: but from him that hath not shall be taken away even that which he hath." The sociologist Robert Merton borrowed this line from the Gospel of Matthew in 1968 to name a mechanism: the Matthew Effect, where a small initial advantage amplifies itself into vast inequality. A famous scientist producing the same result gets more credit; honours flow to those who already have honours. Later came a live experiment with fourteen thousand participants, which for the first time pulled apart how much of success is quality and how much is amplified luck.

The mechanism is not complicated. Merton was watching how credit is allocated in science: the same paper, signed by a famous name, gets read and cited more; citations raise reputation, reputation brings funding and students, the next round starts from a higher place, and round it goes. In 1988 he published a sequel shifting the emphasis from "credit went to the wrong person" toward the general mechanism: cumulative advantage, the rich getting richer at a rate that makes the poor relatively poorer. DiPrete and Eirich's authoritative 2006 review (cited around 3000 times) later established it formally as a general mechanism of inequality, generalising from science to income, education and nearly every stratification phenomenon; Perc's 2014 review systematically counted its traces in empirical data (cited around 700). One sentence for the shape of the mechanism: an advantage is not a one-off asset, it is principal that earns interest.

The paper carries a sharp footnote of its own: it is itself an instance of the Matthew Effect. Much of Merton's empirical material came from Harriet Zuckerman's interviews with Nobel laureates, and he later acknowledged publicly on several occasions that the paper should have been co-authored by Zuckerman and Merton. A senior figure taking sole authorship of collaborative credit demonstrates exactly the mechanism the paper describes. In 1993 the historian of science Rossiter named the gendered version, the Matilda effect (after the 19th-century feminist Matilda Joslyn Gage): women scientists' contributions being systematically transferred to male colleagues. This is not a historical relic; research in 2025 still measured the same-shaped gap in authorship data in communication, political science and sociology. Who gets to accumulate advantage is itself structurally biased.

But the Matthew Effect has a fatal empirical problem: from real-world data alone you cannot tell "an initial advantage got amplified" from "they were simply better to begin with." To tell them apart you would need to run the same world several times and see whether the outcome is the same, and reality only happens once. In 2006 Salganik, Dodds and Watts built actual parallel worlds with an online experiment, which is MusicLab. The reading rule in one sentence: if the same set of songs comes out in the same order in every world, quality decides; if the worlds come out wildly different, social feedback decides.

FIG.01 An artificial music market: quality sets the bounds, not the rank order SIM

The experiment recruited 14,341 real people to listen to 48 songs by unknown bands, with the option to sample and download. The key to the design is the parallel worlds: in one "independent world," listeners could see nothing about anyone else's behaviour and judged by ear alone, which gives a quality baseline for each song; in eight other "social influence worlds," listeners could see each song's current download count, and the eight worlds evolved independently. The second round pushed further, arranging songs by download count from high to low, making the social signal stronger. Two results. First, the stronger the social influence, the more unequal success became: the spread of download counts (measured by the Gini coefficient) widened. Second, and more counterintuitive: the stronger the social influence, the less predictable the outcome, with the same song topping the chart in one world and sinking in another. The authors' summary is worth keeping verbatim: the best songs are rarely at the bottom and the worst songs are rarely on top, but everything in the middle is up for grabs.

The 2008 sequel pushed conditions to the extreme, titled "Leading the Herd Astray," with 12,207 participants. This time the researchers falsified directly: they inverted the chart, displaying the genuinely least popular songs as the most popular. The results have two layers. Looking at most songs, the fake popularity really did fulfil itself, and songs falsely reported as hot really did take off, which is laboratory-grade proof of the self-fulfilling prophecy you learned in R02. But looking at the market as a whole, the inversion did not hold to the end: the best songs recovered their popularity over time. Falsification also had a systemic cost: the correlation between how good a song was and its eventual popularity was reduced, and total downloads in the market fell too, so manipulating the chart does not merely reorder ranks, it degrades the market overall. Put the two experiments together and the precise conclusion is: quality sets the possible range (the best rarely bottom out, the worst rarely win), social influence decides who wins within that range, and it amplifies inequality and unpredictability at the same time.

This set of experiments has special value for the whole red branch. The main evidence in ranking research (R01, on how rankings remake law schools) comes from interviews and observation, and has been criticised as bordering on circular: saying "the ranking changed the schools" with no parallel world to compare against. MusicLab fills exactly that gap: random assignment, parallel worlds, an invertible chart, and the different fates of one song across eight worlds are causal evidence that ordering signals participate in manufacturing rank. It also sets a boundary: the self-fulfilling prophecy operates only within the range quality marks out, and the strong claim that "a ranking can crown any terrible school as number one" is not supported by the experiment.

One open question: "the best songs recover" depends on a recognisable real quality. To what extent is a model's "quality" recognisable and agreed on? If different models' strengths are simply incommensurable, the recovering force may not exist at all. Without that recovering force, chart noise would be permanently locked in. That is the premise to watch most closely when extrapolating MusicLab to AI leaderboards.

The one-line takeaway: quality decides the range you can reach, luck plus feedback decides where in that range you land, and a chart is a machine for amplifying luck into destiny.

Sources / further reading
  • Merton, R. K. (1968). "The Matthew Effect in Science." Science 159(3810):56–63; (1988) "The Matthew Effect in Science, II." Isis 79(4):606–623.
  • Corroboration of his own statements about Zuckerman's co-authorship: Garfield (2004). Scientometrics 60(1); Bacevic (2023). Current Sociology.
  • Rossiter, M. W. (1993). "The Matthew Matilda Effect in Science." Social Studies of Science 23(2):325–341; Goyanes et al. (2025). Scientometrics (the Matilda effect in authorship data).
  • Salganik, Dodds & Watts (2006). Science 311:854–856; Salganik & Watts (2008). SPQ 71(4):338–355 (the inversion experiment); (2009) Topics in Cognitive Science (methods and replication).
  • DiPrete & Eirich (2006). Annual Review of Sociology 32 (the unobserved heterogeneity warning).
  • Further depth in research/deep/D1 §6; simulation in experiments/exp5.
Where to next