REACTOR
K04 Cases Cases · REV.3

SELF-TEST OK · REACTOR v3 · LOADING [ K04 ]…

Markets: Ratings & Fake Reviews

The assembly line behind five-star reviews.

Requires R08 Reflexivity (Soros) Unlocks

What is one star worth? The economist Michael Luca gave a precise answer: each one-star increase in a Yelp rating raises a restaurant's revenue by roughly 5% to 9%. His identification method is regression discontinuity: Yelp rounds the displayed rating, so two restaurants with true averages of 3.24 and 3.26 are almost equally good in reality, but the rounding line splits them apart, showing one as 3 stars and the other as 3.5 stars. The only difference left between the two restaurants is the displayed rating itself, so comparing revenue on either side of the rounding line extracts the causal effect of the rating alone, independent of how good the restaurant actually is. Better still, this effect exists only for independent restaurants, not for chains. A chain already has a national brand backing its reputation, so customers already know what to expect; an independent restaurant has no such brand, so the star rating is the only outside signal a customer can get. This shows that online ratings substitute for exactly what traditional word of mouth used to do. Ratings do not just reflect business, they manufacture business. This is the reflexivity you learned in R08, measured directly in a consumer market. And this is where market ratings differ fundamentally from other fields: the return on gaming can be priced exactly. Once a star has a posted revenue value, fake reviews become an investment product with a defined rate of return. With a rate of return sitting right there, it is only a matter of time before review fraud grows from individual deviance into an industry. From the fake-review supply chain to LIBOR (the financial benchmark rate that was gamed into record fines) to strategic behavior under credit scoring, the three layers are one thing: if a signal has a price, the signal will be manufactured.

The fake-review supply chain: supply, demand and enforcement

Every link in the chain has an enforcement or research record. Scale estimates on-platform: Luca & Zervas found that about 16% of Yelp restaurant reviews are flagged as suspicious by the filtering algorithm, meaning roughly one in six or seven reviews is a suspected fake; and businesses with weak reputations (few reviews, recent bad ones) are more likely to fake, while businesses in more competitive settings are more likely to be targeted with fake negative reviews. Faking is a rational response to circumstances, with the direction set by competitive structure. Off-platform supply: in 2022 Amazon sued the operators of over 10,000 Facebook fake-review broker groups, one of which had over 43,000 members before it was shut down. China's grey market for review fraud has become a complete chain, and according to media accounts of enforcement data, one review-fraud App connected over 36,000 e-commerce sellers with 600,000 fraudsters and manufactured over 200 million fake orders (single source, treated as unsettled evidence). An "employment scale" of six hundred thousand people means this is an industry, not a handful of crooks. The social media metrics counterpart is the Devumi case: the company controlled at least 3.5 million bot accounts and sold over 200 million followers to about 200,000 customers, with at least 55,000 fake accounts stealing real users' names and profile pictures; the New York Attorney General's 2019 settlement was the first to establish that selling fake social media engagement is illegal. The regulatory patch line then straightened out: the FTC (the US Federal Trade Commission) issued its first fine for paid fake reviews in 2019, 12.8 million dollars (a figure of 1280 in the source's ten-thousand-dollar units), settled with Fashion Nova for 4.2 million dollars in 2022 over suppressed negative reviews, and in August 2024 issued the Rule on the Use of Consumer Reviews and Testimonials, comprehensively banning the buying and selling of fake reviews with penalties up to 51,744 dollars per violation. Mechanism tag: outright fraud plus rules arbitrage, with regulation sealing the gaps by legislation after the fact.

A two-way battlefield

Rating manipulation is a two-way weapon, and the same pool of accounts is used to raise your own score and to trash a competitor's. The Bi Zhifei case on Douban has records on both sides: the director sued the platform for "locking in the lowest score" and forcing his film's withdrawal, and the court dismissed the case for insufficient evidence; the same platform is also constantly manipulated in the other direction, with one show receiving 6752 five-star ratings in the six hours before its premiere against only 649 one-star ratings, with review farms quoting prices graded by account quality. Not one episode had finished airing and the reviews were already written; the distribution is its own confession. Rotten Tomatoes (the US film review aggregator) covers both directions too: the PR firm Bunker 15 recruited marginal critics at about 50 dollars a piece to lift scores, taking one film's freshness from 46% (rotten) to 62% (fresh); in the other direction there was the review bombing before Captain Marvel opened, with the audience score pushed down to 33%, after which the platform removed about 54,000 reviews and banned pre-release ratings for good. Bestseller-list arbitrage is more direct: a marketing firm charged about 210,000 dollars under contract to place a book at the top of the New York Times list through coordinated bulk purchasing, and it fell out of the top ten the following week. A "bestseller" that stays on the list for one week is not buying readers, it is buying the advertising value of the list position itself. Mechanism tag: outright fraud and score manipulation, with attackers and defenders buying their ammunition in the same market.

LIBOR: the record-breaking endgame of a self-reported benchmark

Scale "the rating gets gamed" up to financial infrastructure and you get LIBOR. This benchmark rate priced contracts on the order of hundreds of trillions of dollars (the commonly cited figure from the Wheatley Review is around 300 trillion dollars, with mortgages, corporate loans and derivatives all hanging off it), and yet the way it was generated was astonishingly crude: panel banks each day self-reported an estimate of the rate at which they could borrow funds, and the top and bottom were trimmed before averaging. Not based on actual transactions, purely self-reported. The second law of the K01 case library (self-reporting is a breeding ground for fraud) came true here at record scale. Manipulation came in two types: traders nudging their own bank's submission to suit derivative positions, which is opportunistic; and banks collectively lowering submissions during the crisis so that high rates would not expose their own funding difficulties, which is collective survival. In June 2012 Barclays was first to settle with UK and US regulators for about 450 million dollars, and its CEO resigned; UBS, RBS, Deutsche Bank (2.5 billion dollars in 2015) and others settled subsequently, with cumulative industry penalties commonly cited at around 9 billion dollars. Individual criminal liability went back and forth: the trader Tom Hayes became the first individual convicted in 2015 and was initially sentenced to 14 years, and in July 2025 the UK Supreme Court quashed his conviction on the grounds of misdirection of the jury. The institutional endgame is the most telling part: governance reform could not save the benchmark, LIBOR was abolished entirely, the US dollar panel stopped publishing in June 2023, and successor benchmarks (such as SOFR for the dollar) are computed from actual transaction data. It could not be fixed, only the signal source could be replaced: when a measurement is gamed beyond repair, the only defence is to switch to a signal that cannot be self-reported. Mechanism tag: adversarial Goodhart at maximum. This is a course-supplementary case, with facts taken from regulatory penalty notices and the Wheatley Review, not in the research working notes; evidence grading in the sources.

Credit scoring: rational self-modification by the measured

Credit scoring is the mirror image of market ratings: rather than a merchant inflating its own score, the scored party reshapes itself according to published rules, which is the prototype scenario of Y10 strategic classification. Publish the model that predicts repayment probability and an applicant has two routes: one is gaming, stuffing in surface features that fool the model (keywords, temporarily inflating cash flow), which improves the prediction while real creditworthiness is unchanged; the other is improvement, genuinely accumulating savings and reducing debt, so the true value improves too. Hardt et al. formalized this in 2016 as a Stackelberg game (a sequential game where one side publishes rules first and the other responds to them) and proved that a naive classifier degrades sharply under gaming. Later work added two more cuts: distinguishing gaming from improvement is fundamentally a causal inference problem (which features cause the true value and which are merely correlated); and when institutions push the threshold up to the gaming-optimal point to resist gaming, surface-level tricks alone can no longer fool it: to qualify, an applicant now has to make a real change. Making a real change costs something, and the institution does not bear that cost, it falls on the person being classified. But the cost of reshaping a feature is not the same for everyone: disadvantaged groups find it more expensive to change and carry a heavier burden. The bill for resisting gaming ends up with the people least able to game. Institutional-level rating inflation closes this line: before 2008, rating agencies under the issuer-pays structure (the issuer pays to have itself rated) rated as AAA what "on average only BBB supported" (Griffin & Tang's econometrics on 916 CDOs), with an endgame of judicial settlements of 1.375 billion dollars for S&P and 864 million for Moody's. Markets turn every rating into a price, and therefore turn every rating into a target.

One open question: platforms have started mixing non-engagement signals (user surveys, quality proxies) into ranking to hedge against score manipulation, but are these new signals just the next round's target? Faking surveys and manufacturing quality signals is not expensive, and the list of signals that cannot be self-reported looks as though it is getting shorter with use.

The one-line takeaway: once a rating can be exchanged for money, somebody will manufacture ratings; the defence is not moral appeal, it is rewriting the supply and demand curve of faking.

Sources / further reading
  • Luca, M. (2011/2016). "Reviews, Reputation, and Revenue: The Case of Yelp.com." HBS Working Paper 12-016.
  • Luca, M. & Zervas, G. (2016). "Fake It Till You Make It: Reputation, Competition, and Yelp Review Fraud." Management Science.
  • FTC: Cure Encapsulations (2019), Fashion Nova (2022), the Rule on the Use of Consumer Reviews and Testimonials (2024-08); Amazon's announcement of suits against fake-review groups (2022); Devumi: NY AG settlement (2019), FTC (2019).
  • The Douban case: China News Service (2018-08); review farm pricing: Jiemian News. Rotten Tomatoes: Vulture/TheWrap (2023); the bestseller list: Christianity Today (2014, the ResultSource contract).
  • LIBOR (course-supplementary, not in the research working notes): FSA/CFTC/DOJ penalty notices against Barclays (2012); Wheatley Review (2012); the UK Supreme Court's ruling in the Hayes case (2025-07).
  • Hardt, M., Megiddo, N., Papadimitriou, C. & Wootters, M. (2016). "Strategic Classification." ITCS; Miller, Milli & Hardt (2020), ICML; Milli, S. et al. (2019). "The Social Cost of Strategic Classification." FAT*.
  • Griffin, J. & Tang, D. (2012). "Did Subjectivity Play a Role in CDO Credit Ratings?" Journal of Finance; DOJ settlement announcements with S&P (2015) and Moody's (2017).
  • Working notes: research/04-cases-across-domains.md §4A, §6–7, §9; research/05-ai-evals-reactivity.md (strategic classification); research/deep/D6 sub-thread 3; research/deep/D3 sub-threads 5–6.
Where to next