Meta’s retrieval engine reduces tens of millions of eligible ads to a few thousand candidates before any auction logic runs. That stage is called Andromeda, it has been core infrastructure since 2025, and Meta reports it delivering a 6% recall improvement to the retrieval system and an 8% ads quality improvement on selected segments. Reading the engineering write-up changed how I structure accounts more than any product announcement in the last three years.
For roughly a decade the craft on Meta was audience construction. Stacked interests, lookalike percentages, exclusion logic, layered demographics. It was genuinely skilled work, and I was good at it. It is now close to worthless, and the reason is architectural rather than fashionable.
The lever moved upstream
Retrieval happens before ranking. If a system is selecting a few thousand ads out of tens of millions by predicting per-user engagement, then the question it is answering is not who an advertiser wants to reach. It is which piece of inventory is most likely to work for the person currently loading the app. An audience definition is a constraint on that search. Constraints on a search that good mostly remove options it would have found anyway.
The capacity figures make the asymmetry concrete. Meta describes a 10,000x increase in model capacity for personalisation, a 100x improvement in feature extraction latency and throughput, over 3x on end-to-end inference queries per second, and a further 10x from model elasticity. Whatever a media buyer contributes by hand-picking three interest clusters is not competing with that. It is competing for the privilege of narrowing it.
Meta reports a 22% increase in ROAS for advertisers who turned on AI-driven targeting.
The layer above tells the same story. GEM, the ranking model, is described as 4x more efficient at driving ad performance gains for a given amount of data and compute than the ranking models it replaced. Retrieval got better, ranking got better, and neither improvement is available to an advertiser who keeps the system on a short leash.
What actually remains under an advertiser’s control
If targeting is no longer the lever, three things are, and all of them are inputs to the model rather than instructions to it.
- Creative volume and variance. A retrieval engine choosing between candidates needs candidates. Meta says more than a million advertisers produced over 15 million ads in a month with its generative tools, and reports a 7% lift in conversions for businesses using image generation. The scarce resource is distinct creative, not clever segments.
- Conversion signal quality. Everything upstream is trained on what gets sent back. A broken or partial signal does not produce cautious optimisation, it produces confident optimisation toward the wrong event.
- Catalogue and feed hygiene. For any commerce account the feed is the ad inventory. Bad titles and missing attributes remove items from consideration before a bid is ever calculated.
What this looked like on a real account
On one e-commerce account I run, the structure had 14 ad sets built around interest stacks that had worked well in 2023. Performance had been slowly degrading for months in a way that looked like creative fatigue. It was not. Consolidating to a small number of broadly-targeted campaigns and moving the freed time into producing more distinct creative reversed the trend, and the campaigns exited the learning phase faster because each one was finally getting enough events to learn from.
The uncomfortable part was that the 14 ad sets represented craft. Someone had thought carefully about each one. The system simply had a better view of the same question, and the segmentation was splitting the signal it needed into fourteen pieces too small to be useful.
The disclosure question nobody has settled
There is a second-order problem arriving. The IAB has updated its guidance toward targeted disclosure for specific uses such as AI-generated images and video, rather than a blanket label on anything AI-touched. That is the sensible position, but it lands awkwardly when the generative tooling is inside the ad platform and producing assets at the volume above. An advertiser can end up shipping machine-made creative without ever making a deliberate decision to do so.
The one number here that is not Meta’s
Independent evidence is thinner than it should be, and what exists measures adoption rather than performance. An EMARKETER survey found 57% of US digital ad buyers planning to use AI to assist the buying process through automated tools such as Advantage+ and Performance Max. That tells you where the industry went. It does not tell you the industry was right, and the same coverage describes Advantage+ as having produced mixed results, which is a phrase that never appears in a vendor case study.
Where I might be wrong, and it is a real risk
Every performance number in this piece comes from Meta, about Meta’s own automation, published to encourage advertisers to hand over more control. The 8% quality figure is explicitly scoped to selected segments, which is the kind of qualifier that does a lot of quiet work. No independent replication exists, and none is possible from outside, because no third party can run the counterfactual.
So the honest version of the claim is narrower than the title. I do not know that Andromeda is as good as Meta says. What I observe is that hand-built audiences stopped outperforming broad targeting in my own accounts, which is a much weaker statement, and consistent with a duller explanation: Meta may have degraded manual targeting’s effectiveness rather than improved automation’s. From an advertiser’s chair those are indistinguishable, and the correct response happens to be identical either way.
The place I would expect to be proven wrong is narrow, high-consideration B2B, where the addressable universe is small enough that a constraint really does carry information the model cannot infer from behaviour. I have not tested that carefully enough to argue it, and anyone claiming a universal answer on an account structure question is selling something.
What I no longer believe is that audience construction is where the skill lives. The skill moved to what gets fed in: enough distinct creative to give retrieval something to choose from, a conversion signal that is actually true, and a catalogue that does not disqualify itself. That is less satisfying than building a clever exclusion stack. It is also what the architecture rewards.