Why I stopped trusting platform-reported ROAS

Why I stopped trusting platform-reported ROAS

Every ROAS figure inside an ad platform’s interface was calculated by the company that gets paid when that figure looks good. That is not a conspiracy. It is a conflict of interest built into the architecture, and it is why I stopped treating platform-reported ROAS as a measurement. Across twenty-two years of buying media and more than $8M in lifetime managed ad spend, that number has kept a permanent job in my workflow and lost all standing as evidence.

The sharpest illustration is not a critic’s paper. It is Google’s own research, and the useful half is the correction it published nine months after the headline.

The platform published its own grade

In July 2011, Google Research published its Search Ads Pause studies: over 400 studies on paused accounts, watching organic click volume with the ads switched off. The finding, in Google’s own words: on average, the incremental ad clicks percentage across verticals is 89%. Read plainly, 89% of the traffic search ads generate is not replaced by organic clicks when the ads stop.

I watched that number travel for a decade: pitch decks, quarterly reviews, arguments with finance directors. A real result, published by the party with most to gain from it, and used as though it held everywhere.

Then Google published the number that kills the shortcut

In March 2012, Google Research released a meta-analysis of 390 of those same pause studies, split by where the advertiser already ranked in organic search:

  • 50% of ad clicks are incremental when the advertiser holds the top organic result.
  • 82% are incremental when the organic result sits between ranks 2 and 4.
  • 96% are incremental when the organic result ranks lower than 4.

50% of the ad clicks that occur with a top rank organic result are incremental. 96% of the ad clicks are incremental when the advertiser’s organic result ranked lower than 4. — Google Research, meta-analysis of 390 Search Ads Pause studies, 2012

Half the value, or nearly all of it, decided by a variable that lives outside the ads account entirely. Same platform, same ad format, same method, same researchers. There is no honest way to read that table and still believe a single correction factor exists.

Google's own studies: how incremental a paid click is, by the advertiser's organic rank

Google's own studies: how incremental a paid click is, by the advertiser's organic rankHorizontal bar chart showing that 50 percent of paid clicks are incremental when the advertiser already holds the top organic result, 82 percent when its organic result ranks between two and four, and 96 percent when it ranks below four.Advertiser holds top organic result50%Organic result ranks 2 to 482%Organic result ranks below 496%

The platform's own experiments put the value of a click on a slider, which is why one blanket discount cannot be correct for two different accounts. Source: Google Research, 2012.

Folk multipliers are the same mistake in costume

So the trade invented shortcuts. Halve Meta’s reported ROAS. Take a third off Google. Treat every view-through conversion as fiction. Divide by 1.4 because a podcast said so. I have written before about the gap between attribution and incrementality; this habit is narrower and worse. A folk multiplier feels like sophistication and is the original error in costume: a fixed number applied to a quantity that is not fixed. The advertiser has swapped the platform’s biased estimate for a personal one with nothing underneath it.

One account where the error ran in both directions

A lead-gen client of mine runs two things out of one ad account: cold prospecting across a set of metro areas, and a retargeting pool fed by site visitors. Reported ROAS ranked them the way it nearly always does, retargeting far out in front. I sorted the metros into matched pairs on pre-period lead volume and went dark in one market of each pair for six weeks.

The retargeting number collapsed, the textbook result: most of those people were in the funnel already and arrived anyway. The prospecting campaign did the opposite. Lead volume in the dark markets fell by more than the platform had ever claimed credit for. That business takes a large share of its enquiries by phone, days after first exposure, and none of those calls were stitched back to a click. The platform was undercounting the campaign it ranked last and overcounting the one it ranked first.

Now apply a flat discount to that account. Halving both numbers preserves the ranking, and the ranking was the error. The tidy correction would have starved the only campaign creating demand.

The error is large and unstable, not constant

The most complete public accounting comes out of Meta. In Marketing Science in 2023, Brett Gordon, Robert Moakler and Florian Zettelmeyer worked through 663 large-scale experiments at Facebook, asking whether standard non-experimental methods could recover what the randomised tests had already found. Against true experimental lifts of 29%, 18% and 5% for upper, middle and lower funnel outcomes, stratified propensity score matching returned median lifts of 173%, 176% and 64%. Double machine learning, the better of the two, still returned 83%, 58% and 24%.

The ratios matter more than the levels. At the lower funnel the propensity-matched estimate overshot by roughly thirteen times; at the upper funnel, by six. Same platform, same method, same data, and the size of the error swung by a factor of two on funnel position alone. An earlier paper from the same group, on 15 US advertising experiments at Facebook covering 500 million user-experiment observations and 1.6 billion ad impressions, reached the same verdict.

Lower-funnel lift at Facebook: what the experiment found against what the models reported

Lower-funnel lift at Facebook: what the experiment found against what the models reportedBar chart comparing a 5 percent lower-funnel lift measured by randomised experiment against 24 percent estimated by double machine learning and 64 percent by propensity score matching.5%24%64%RandomisedexperimentDouble machinelearningPropensity scorematching

Both non-experimental methods pointed in the right direction and landed nowhere near the truth, which is the practical trouble with correcting a reported number by feel. Source: Gordon, Moakler & Zettelmeyer, Marketing Science, 2023.

Nor is this academic. Seer Interactive analysed $1.05 million of April 2025 Meta spend after Meta shipped its own incrementality setting. Meta put non-incremental conversions at 13%; GA4, on the same spend, put them at 33%. A twenty-point gap on the only question that matters, and one of the two graders is selling the media.

Measure your own gap instead

The only defensible correction factor is one measured on the account it gets applied to. That means a holdout, and holdouts are less exotic than they sound.

  1. Choose a unit you can genuinely switch off. Geography is usually the only one: postcode, DMA or country.
  2. Split on pre-period outcomes, not the alphabet. Matched pairs beat a random draw when units are few.
  3. Switch it off rather than trim it. A budget cut changes what the algorithm buys, not only how much.
  4. Run past the longest attribution window plus the real sales-cycle lag. Six weeks is a floor for lead-gen.
  5. Compare outcomes from your own source of truth. The whole point is a number the platform did not produce.
  6. Record incremental result over reported result. That ratio is the discount factor, and it belongs to that campaign, that offer, that season.
  7. Run it again next quarter. It moves.

Ordinary work, all of it. Most accounts skip it not from difficulty but because the answer might justify a smaller budget.

The part I could be wrong about

There is a hard statistical limit here, from Randall Lewis and Justin Rao in the Quarterly Journal of Economics. Across twenty-five large field experiments with major US retailers and brokerages, collectively representing $2.8 million in digital advertising expenditure, the median confidence interval on return on investment is over 100 percentage points wide. Individual sales are so volatile against the per-capita cost of advertising, with a coefficient of variation of 10 being common, that informative experiments can easily require more than 10 million person-weeks.

So here is a position I hold and could be wrong about. Below roughly $50,000 a month, campaign-level holdouts are not worth running. At that spend the test cannot separate a ROAS of 1 from a ROAS of 4, and a noisy answer held with confidence does more damage than a biased one held with suspicion. What I tell those accounts instead: test at channel level once or twice a year, and spend the rest of the effort making the platform’s signal less wrong rather than auditing it. Server-side tagging, offline conversion imports, deduplicated events, real order value on the event.

The objection is obvious. If the honest advice to a small advertiser is that the answer cannot be measured, they go on trusting the number I just told them to distrust. That is a poor place to leave someone. I have not found better, and if anybody shows me a holdout design that reads reliably at $20,000 a month, I will change my position that week.

Platform-reported ROAS is not worthless and not a lie. It is a house number, computed by the house, and it ranks creatives inside one auction perfectly well. It has no business deciding how much of a client’s money that auction deserves. Google’s own researchers put a paid click anywhere between half and nearly all of what the platform counts, on facts the platform cannot see. One multiplier over that range is guessing with extra steps.

Sources

  1. Google Research, 2011
  2. Google Research, 2012
  3. Gordon, Moakler & Zettelmeyer, Marketing Science, 2023
  4. Gordon, Zettelmeyer, Bhargava & Chapsky, Marketing Science, 2019
  5. Lewis & Rao, Quarterly Journal of Economics, 2015