- Home
- The Thinking Wire
- 1.28%: How Much of the Ranked Result Set Survives Into the AI Answer
1.28%: How Much of the Ranked Result Set Survives Into the AI Answer
Only 1.28% of the products ranking in traditional Google search also appear in AI Mode for the same search on the same day. That figure comes from Productrise, which tracked more than 2 million product listings across more than 100,000 SERPs and AI Mode responses over 23 days (August 9 to 31, 2026), in the US and the UK, matching products on Google’s stable product identifier and running both surfaces simultaneously on the same calendar day.
Productrise sells e-commerce analytics, and this is their own analytics blog. The methodology is theirs, the sample is theirs, and we have seen no independent replication of it. Read the numbers as one vendor’s large measurement rather than as settled fact.
Even discounted for that, 1.28% is not a ranking story. Position 1 versus position 8 is a question about order. This is a question about existence.
The Shape of the Filter
Traditional search shows 27.8 products per query. AI Mode shows 3.9. The overlap between the two sets is 0.94 products. Less than one product, on average, appears on both surfaces for the same query on the same day.
That is a different operation from ranking. A ranked list is a full inventory with an opinion attached, and the opinion is inspectable: you scroll past the top result, you see what came second and eleventh, you form your own view of whether the ordering makes sense. AI Mode returns about four products and no list. The roughly 27 that did not make it leave no trace in the answer.
Productrise’s own July 2026 analysis found AI Mode returning approximately 95% fewer products than traditional search. The August run puts a sharper number on the same behavior and adds the part that matters commercially.
The Survivors Are More Expensive
For matched products (the ones ranking in both surfaces), AI Mode is 21.6% more expensive.
The mechanism sits in the seller layer. Prices differ between the two surfaces 38.1% of the time. When they differ, AI Mode is the dearer of the two 68.4% of the time, with a median difference of 22.2%. The main seller differs on 49.6% of matched products, which means that on roughly half the products present in both places, the buyer is being routed to a different merchant entirely.
The same physical product, the same query, the same day, a different seller, and more often than not a higher price. Google’s ranked results and Google’s generated answer disagree about who should sell you the thing.
Katelyn Geary, Senior SEO Strategist at Break The Web, framed the consumer side of this in the Productrise piece: “AI search was supposed to eliminate the friction of comparing prices across five different tabs, but if it only serves pricier inventory by default, it hasn’t saved us any work.” The friction did not disappear. It moved out of view, which is a worse place for it to live than five open tabs.
Callum Lockwood, Director of Organic Search at Re:signal, pointed at the merchant side in the same article: “The seller swap on almost half of matched products is the bit that should worry retailers more than the price gap itself.” A price delta is a margin problem you can model. Being substituted out of the answer on half your matched catalogue is a distribution problem, and it is invisible from inside your own analytics, because the impression you never received generates no data.
Why This Is a Governance Question and Not a Marketing One
Marketing teams will read these numbers as an optimization brief. That reading is available and mostly beside the point.
Ranking bias in search is a problem regulators already recognize and already have a vocabulary for, because a ranked list is auditable. You can pull the SERP, count positions, compare the treatment of first-party and third-party inventory, and argue about it with evidence in hand. That recognition rests on the list being there to inspect.
A generative answer removes the list. There is no position 8 to point at, no comparison set, no visible indication that a cheaper listing for the same product existed somewhere in the roughly 27 that were dropped. The output is about four products and a paragraph of confident prose. The selection logic that produced them is not exposed, and the counterfactual is not recoverable from the answer itself. Someone had to run a 23-day parallel measurement to see it at all.
That is the structural change. The audit surface for recommendation integrity disappeared at the exact moment the recommendation started carrying more weight with the buyer. We wrote about the missing standards layer underneath this in agentic commerce governance, and about how citation concentration reshapes visibility in AEO’s four separate games. This is the same problem measured on the axis that ends up on someone’s invoice.
The Same Blind Spot Inside Your Own Systems
Any team shipping an LLM-backed recommendation, whether it surfaces products, vendors, candidates, or internal documents, has built the identical structure. A retrieval step produces a candidate set. A generation step selects a few and writes prose about them. The user sees the prose.
Ask what you can currently prove about your own pipeline:
- What fraction of the retrieved candidate set reaches the user’s answer? If nobody has measured it, you do not know whether your system narrows sensibly or arbitrarily.
- Do the selected items differ systematically from the ones dropped, on any attribute that carries money? Price, margin, vendor, contract status, recency. Productrise found a 21.6% price skew on matched products, on a surface where nothing exposes it.
- If a user asks why one option appeared and another did not, can you answer? Not with an explanation the model generates after the fact. With a logged candidate set and a recorded selection.
We have written before about scoreboards broken at both ends. This is a specific instance: the input set is unlogged and the output is unfalsifiable, so the metric everyone watches is engagement with an answer nobody can check.
Do This Now
Pick your highest-traffic LLM-backed recommendation surface and log the candidate set alongside the final answer for a full billing period. Just the two sets, timestamped and joined by request ID. No new infrastructure, no model changes.
Then compute three numbers: what fraction of candidates survive into the answer, whether the surviving items skew on price (or margin, or whichever attribute carries commercial weight in your context), and how often the top-ranked candidate is not the one presented.
If the survival rate is low and the skew is real, you have a recommendation integrity problem, and you now have the evidence to fix it before a customer or a regulator finds it first. If the numbers come back clean, you have a baseline and a monitor. Either outcome beats shipping a filter nobody has ever looked through.
This analysis synthesizes Google AI Mode prefers more expensive products (Productrise, September 2026).
Victorino Group helps teams instrument LLM-backed recommendation systems so the selection behind an answer can be inspected and defended. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation