DevXen
All insightsLocal growth5 minute read

Reviews and Revenue: What the Evidence Shows

A causal study found a one-star rating increase lifted revenue 5 to 9 percent, but only for independents. Chains saw no effect at all.

"Reviews matter" is not a controversial claim, which is why it is usually asserted rather than evidenced. It is also one of the rare marketing questions with a properly identified causal estimate behind it, and the details of that estimate change what you should do about it.

The study

Michael Luca, "Reviews, Reputation, and Revenue: The Case of Yelp.com," Harvard Business School Working Paper 12-016.

Luca combined Yelp review data with restaurant revenue data obtained from the Washington State Department of Revenue, covering every restaurant in Seattle that reported revenue between January 2003 and October 2009. That is 3,582 restaurants, around 1,587 open in a given quarter, of which 143 were chain affiliated.

Tax-reported revenue, an entire city, and a period spanning both before and after Yelp arrived.

The identification problem, and how he solved it

The obvious analysis is worthless on its own. Good restaurants get good ratings and good revenue, so a correlation tells you nothing about causation.

Luca exploited a quirk of how Yelp displayed ratings. The average was rounded to the nearest half star. A restaurant at 3.24 displayed as 3 stars. A restaurant at 3.25 displayed as 3.5 stars. Those two restaurants are essentially identical in quality, but consumers saw a half-star difference.

Comparing restaurants within 0.1 stars of a rounding threshold isolates a change in displayed rating that is unrelated to actual quality. This is a regression discontinuity design, and it is what turns the finding from a correlation into a causal estimate.

He also ran a McCrary density test to check whether restaurants were gaming ratings just above thresholds. They were not clustering there, so manipulation does not explain the result.

What he found

A one-star increase in rating caused a 5% to 9% increase in revenue. The main regression discontinuity estimate was 0.094 with a standard error of 0.041, significant at the 5% level, across 2,169 observations and 854 restaurants. Widening the bandwidth gave 0.054.

The effect came entirely from independent restaurants. For independents, the rating coefficient was 0.065 and highly significant. For chains, it was 0.005 with a standard error of 0.025, statistically indistinguishable from zero.

Reviews substituted for chain brand. As Yelp penetration increased in a market, revenue share shifted away from chains and toward independents.

Not all reviews carried equal weight. Ratings backed by more reviews produced larger responses, consistent with consumers treating a rating with more data behind it as a more reliable signal. Reviews from Yelp "elite" reviewers had nearly double the impact of others. The size of a reviewer's friend network had no effect at all.

For scale, Luca compares his estimate to Jin and Leslie (2003), who found a restaurant hygiene grade moving from B to A produced a 5% revenue increase.

What it does not establish

The chain finding is the one that gets dropped from every summary. If you are a franchise or a known national brand, this study found no revenue effect from ratings. The mechanism is information: reviews matter most where the customer has no other way to judge quality. That is precisely the position of an independent local business, which is why the finding is useful to you and irrelevant to a national chain.

Other limits worth stating.

It is Seattle restaurants, 2003 to 2009. Extending it to a mobile mechanic in Wollongong in 2026 is an inference, not a result.

It is Yelp, which is not the dominant review platform in Australia. Google reviews now occupy that position, with different display, different rounding, and different search integration.

Luca notes explicitly that there is no hard evidence in the paper on the correlation between Yelp ratings and objective quality. The study shows ratings move revenue. It does not show ratings measure quality.

And the effect is identified at the rounding threshold. It measures what happens when the displayed rating jumps, which makes it a strong estimate for businesses near a threshold and a weaker guide elsewhere.

What it justifies

Systematically asking for reviews is a defensible spend. For an independent local business, this is one of the better-evidenced marketing activities available, and asking costs almost nothing. Automating the request after a completed job is an appropriate use of automation.

Volume has value beyond the average. Because response was stronger for ratings with more reviews behind them, a 4.5 average from 200 reviews is worth more than a 4.5 from 6. This matters for a business that has been trading for years with a handful of ratings.

Rounding thresholds are worth knowing about. The entire effect in this study was identified on the displayed rating, not the true one. If you sit just below a threshold, a small number of additional positive reviews is disproportionately valuable. That is not a trick, it is what the data showed about how consumers actually read ratings.

Do not buy reviews. Beyond being illegal under Australian Consumer Law and a breach of every platform's terms, the paper's own gaming analysis makes the economic argument: if manipulation were widespread and effective, the underlying effect would not be measurable. The value of the signal depends on it being real.

The claim this evidence supports is narrow and strong: for an independent local business, a better displayed rating causes more revenue, and the effect is large enough to justify making review requests a standard part of finishing a job.

Sources

  • Luca, M. "Reviews, Reputation, and Revenue: The Case of Yelp.com." Harvard Business School Working Paper 12-016.
  • Jin, G., and Leslie, P. "The Effect of Information on Product Quality: Evidence from Restaurant Hygiene Grade Cards." Quarterly Journal of Economics, Vol. 118, No. 2, 2003.
  • McCrary, J. "Manipulation of the Running Variable in the Regression Discontinuity Design: A Density Test." Journal of Econometrics, Vol. 142, No. 2, 2008.