Six in Ten AI Citations Point Below Rank 100,000. Three Linked Sites Explain Part of It.

TV
Thiago Victorino
8 min read
Six in Ten AI Citations Point Below Rank 100,000. Three Linked Sites Explain Part of It.

On 2 September 2026, Trellner Research put 380 buyer-intent categories to two Perplexity search models through OpenRouter, 760 calls in all, and kept every source the answers cited. The haul was 7,534 citations spanning 2,055 distinct domains. Each domain was then looked up in the Tranco daily list for 2026-09-01.

59.8% of those citations point at domains ranked worse than #100,000. 23.4% point at domains that are not in the top million at all. The median Tranco rank of the 5,768 citations that do land on a ranked domain is 71,611.

That is the corpus. Not the authoritative web, not the trade press, not the analyst firms. A long tail with a median rank somewhere past seventy thousand.

The Practice Rests on Two Assumptions Nobody Measured

Every AI-visibility engagement I have seen priced in the last year quietly assumes two things. First, that the pool of documents an answer engine draws from approximates the credible web, so being cited means being credible. Second, that citation share behaves like a metric: a number you can report at a level, compare across quarters, and benchmark against a competitor.

Trellner’s report (TR-2026-009, no individual byline, published under Trellner Research) takes the first assumption apart. The Landmark Brief takes the second apart in the same week. Neither piece needed the other to land, and together they invalidate most of what gets sold as answer-engine optimization.

215,128 Buying Guides, Six Blog Posts Each

The distribution alone would be an interesting finding. What makes it a governance problem is what sits inside the tail.

Trellner identified three domains in the cited set that share Cloudflare nameservers and were registered between December 2023 and May 2024: worldmetrics.org, gitnux.org and wifitalents.com. The report describes them as under apparently common control, which is the honest way to put it. Shared infrastructure and adjacent registration dates are strong circumstantial signals. They are not a registry record of ownership.

Between them, those three sites have published 215,128 machine-generated best <category> pages. Against six blog posts each.

That ratio is the tell. A publisher with an editorial operation produces articles at a rate a human newsroom can sustain. A publisher producing two hundred thousand comparison pages against a handful of posts is not running an editorial operation. It looks like a fill operation aimed at the retrieval surface answer engines expose.

The top of the citation table is more mixed, and worth reading carefully:

  • g2.com, 291 citations (3.86%), Tranco rank 4,027
  • reddit.com, 261 citations (3.46%), Tranco rank 105
  • guideflow.com, 194 citations (2.57%), Tranco rank 177,039
  • gartner.com, 158 citations (2.10%), Tranco rank 1,766

Three of the four are what you would expect. The third is a vendor’s own marketing blog, ranked 177,039, cited more often than Gartner. The report does not say how it got there.

I have argued before that adversaries can poison AI buyer answers and that the durable play is governing what AI trusts rather than chasing AEO tactics. Both were arguments from anecdote. This is the first quantification of the underlying corpus I have seen with a published method, a stated model set, a dated Tranco snapshot, and dataset plus scripts released under CC BY 4.0. Anyone can rerun it and disagree with the number.

The KPI Moves Threefold on Platform Mix Alone

Jim Cook’s brief works the other end of the same problem. He examined a June 2026 McKinsey study covering 2.6 million citations, 25 consumer packaged goods brands, 89,119 unique domains, UK and US, from October 2025 to May 2026.

The headline metric in that study is brand-owned citation share. Across models, it ranges from three percent to ten percent. Cook’s point: that is a threefold spread produced by nothing except platform mix. Change which engines are in your sample, and the number moves by a factor of three without a single thing changing about the brand, its content, or its authority.

His formulation:

A number that moves threefold on platform mix alone, and five-fold across published studies, can support a trend line inside a single frozen instrument. It cannot support a level, a benchmark, or a comparison.

Cook also flags something that belongs in every discussion of this data: the vendor that supplied the citation data for that June 2026 study also sells generative engine optimization monitoring solutions. The party measuring the problem sells the remedy the measurement implies. That does not make the data wrong. It makes independent replication a requirement rather than a nicety, and it is one more reason Trellner’s open dataset matters more than its headline percentage.

The Corpus Moves Under You Without Notice

The last piece is timing. State of Brand reported that after ChatGPT’s August 6, 2026 change, listicle citations fell from 15.77% to 7.80%, and product pages became 16.39% of retrieved pages. That publication discloses no sample sizes for its own percentages, so treat the magnitudes as directional. The direction is what matters here.

An agency whose programme was built on listicle placements had its retrieval surface cut roughly in half by a change it did not see coming and could not have influenced. I wrote about that specific reallocation in the ChatGPT model switch and what it does to brand function. Put it next to Cook’s threefold spread and the picture is complete: the instrument is unstable, and the corpus it reads is unstable, and the two instabilities are independent.

What To Do This Week

Stop reporting citation share as a level. Report it as a trend inside a frozen instrument, and freeze the instrument explicitly: fixed model list, fixed prompt set, fixed cadence, version-stamped. When a platform changes, the series breaks. Say so in the report rather than letting the line continue as if nothing happened.

Then run a source-provenance audit on the citations you already have. This is the control you actually own, and it takes an afternoon:

  1. Export every domain cited in your category over your last full reporting period, from whatever monitoring you already pay for.
  2. Look each one up in the Tranco daily list. It is free, dated, and reproducible. Record the rank and the date of the list you used.
  3. For every domain below rank 100,000, check two things: how many pages it has indexed in your category shape (best <category>, comparison, alternatives), and how much non-templated content it has published. A large first number against a near-zero second number is the fill-operation signature.
  4. Check registration date and nameservers across the domains that keep appearing together. Shared infrastructure across recently registered sites is a cluster signal worth flagging, and worth stating as a signal rather than as proven ownership.
  5. Write the result as a named list your marketing and legal teams both hold: sources we consider credible, sources we consider manufactured, sources we have not yet classified.

That list is a durable asset. It survives the next model change, the next platform reshuffle, and the next vendor benchmark. Your citation share does not.

The teams that come out of the next two years with defensible AI visibility will be the ones who can name, domain by domain, what the machine read before it recommended them.


This analysis synthesizes Manufactured Sources Behind AI Recommendations (Trellner Research, September 2026), Citation Share Is Not a Metric Yet (Agentic Landmark, The Landmark Brief, September 2026), and The Pages Your Agency Built to Win AI Citations Just Lost Half of Them (State of Brand, August 2026).

Victorino Group helps companies build source-provenance controls for the corpus answer engines read about them. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation