· Updated · 18 min read · Geoptimizer Team

AI Citation Sources: The 68% Concentration Problem

  • generative-engine-optimization
  • ai-citations
  • digital-pr
  • citation-sources
  • ai-visibility
AI Citation Sources: The 68% Concentration Problem

The widely quoted figure — that the top 15 domains capture 68% of all consolidated AI citation share — comes from 5W Public Relations' AI Platform Citation Source Index 2026, a synthesis of six published studies covering more than 680 million citations. Independently measured data broadly supports the direction but not the precision: by our sum of Similarweb's published US shares for January–February 2026, ChatGPT's top 15 domains account for 56.7% of its citations — while Google AI Mode's top 15 account for just 38.2%. And of ChatGPT's concentrated share, only about 37 points sit on domains a content or PR team can realistically earn a place on. That gap between "concentrated" and "addressable" is where your next quarter's budget should be decided.

If you own editorial partnerships, guest content or community programs, this is the number your plan quietly depends on. Pitch against the wrong list and you spend a quarter earning placements the engines never retrieve.

Where the 68% figure actually comes from

Credit where it's due: 5W did something nobody else had bothered to do, which is line up six independent citation studies and produce a single ranked board of the domains that matter. The index draws on "more than 680 million individual citations across ChatGPT, Google AI Overviews, Perplexity, Gemini, and Claude, drawn from six of the largest published citation studies conducted between August 2024 and April 2026," with Profound, Semrush, Ahrefs, Search Engine Land, Muck Rack and Visual Capitalist named as contributing research. Its top 15, in the index's own order, reads: Reddit, Wikipedia, YouTube, LinkedIn, Forbes, Amazon, Business Insider, TechRadar, Reuters, The New York Times, Financial Times, Time, Axios, Quora and The Guardian.

"For 25 years, the most consequential algorithm in communications was Google PageRank," 5W founder Ronn Torossian said in the release. "That era is over."

It is a useful board. It is not a measurement. The index publishes no methodology for how those six studies were selected, normalized or weighted into a single 68% — and the studies themselves span different engines, countries, query mixes and a twenty-month window of extremely fast platform change. 5W also sells advisory work in this category, which doesn't make the number wrong but does mean it should carry an attribution rather than be asserted as fact. Treat 68% like any directional industry index: good for pointing, bad for arithmetic.

So what happens when you use only numbers somebody actually measured?

What the independently measured data says

The most useful public table comes from Similarweb, which published citation shares for ChatGPT (web browsing) and Google AI Mode across US queries in January–February 2026. Sum its published percentages and the picture sharpens fast. ChatGPT's top two domains alone — Wikipedia at 13.15% and Reddit at 11.97% — take 25.1% of everything cited. The top 10 take 47.9%, the top 15 take 56.7%, and the top 20 take 63.6%. (Those sums are our arithmetic on Similarweb's published shares, not Similarweb's own claim.)

That's a striking result on its own: ten domains supply nearly half of everything ChatGPT cites. And it holds up under a second, independent measurement. Ahrefs' Brand Radar, sampling broad US queries five months later in July 2026, puts ChatGPT's combined top-10 share at 49.5%, led by Reddit at 16.7% and Wikipedia at 8.9%. Two different tools, two different windows, two different query samples, and they land within two points of each other. When measurements that share no methodology agree that closely, the underlying finding is usually real.

They disagree about almost everything else, though. Ahrefs' positions three through ten include merriam-webster.com, engineerfix.com, legalclarity.org and dictionary.cambridge.org; Similarweb's equivalent slots are openai.com, walmart.com, YouTube and LinkedIn. The shape of concentration is stable; the membership below the top two is not. Any plan built on positions 3–15 of a single published list is building on the noisiest part of the data.

The bigger crack in "68% across engines" is that it averages away the engines. Google AI Mode's top 15, from the same Similarweb table, reach only 38.2% — nearly 20 points less concentrated than ChatGPT's, with fandom.com rather than Wikipedia at the top. The academic anchor points the same way: an arXiv analysis of 366,087 citations from 12 AI search models found OpenAI models the most concentrated (Gini 0.83) and Google the least (0.69), with Perplexity citing 1,430 unique news sources against OpenAI's 707. One caveat secondary coverage keeps dropping: that paper's "top 20 sources = 67.3%" figure applies to news citations, which were only 9% of its dataset.

Scope is where this category trips people up most. You'll see "Wikipedia accounts for 47.9% of ChatGPT citations" repeated in decks everywhere. Profound's original says Wikipedia is 47.9% of ChatGPT's top-10 slice; its share of total citation volume is 7.8%, and Similarweb's independent read puts it at 13.15%. A six-fold overstatement is the kind of thing that gets a whole strategy memo discounted, so check the denominator before it reaches a client deck.

The 37% rule: how much concentration you can actually earn

Here's the part that changes budgets. Concentration only matters to you if the concentrated domains are ones you can get onto — and a large slice of them are not.

Walk Similarweb's ChatGPT top 20 and sort it by whether a content or PR team could plausibly earn presence there. The addressable column: Wikipedia (13.15%), Reddit (11.97%), YouTube (2.67%), LinkedIn (2.42%), Reuters (2.27%), Facebook (1.76%), GitHub (1.62%) and Forbes (1.38%). That sums to roughly 37.2% of ChatGPT's citations — against a top-20 total of 63.6%.

The other 26 points are not a media plan. They're openai.com (6.21%), the model's own domain; walmart.com (2.90%), media-amazon.com (1.94%), ebay.com (1.75%) and amazon.com (1.71%), marketplaces you can only enter if you sell physical products; nih.gov (2.22%), government research; google.com (2.17%), apple.com (1.48%) and yahoo.com (1.44%); fandom.com (1.29%); and squarespace-cdn.com (1.29%), which is a content delivery network. No pitch, no partnership and no thought-leadership program touches those.

So the honest version of the headline is that the citation pool is genuinely concentrated, and roughly three fifths of that concentration is addressable. Still an enormous prize — a third of one engine's entire citation volume sitting behind eight domains — but it reframes the work. You are not "getting into the top 15." You are getting into the six or seven that accept outsiders.

Apply one further discount, because even those eight aren't equally open. Wikipedia is a lagging indicator, not a lever: it requires notability established by significant independent coverage, and its conflict-of-interest rules restrict brands and paid editors from editing their own articles without disclosure. You earn a Wikipedia presence by earning everything else first. So the practical, this-quarter list is shorter still — Reddit, LinkedIn, YouTube, one or two earned-editorial outlets, and whatever your category's own board looks like.

That last clause is doing more work than it looks.

Your top 15 is not the top 15

Category-level data demolishes the idea of one universal source list. Attrifast ran 1,200 buyer-intent prompts three times each across ChatGPT Search, Claude, Gemini and Perplexity between April 12 and May 14, 2026 — 14,400 prompt-engine runs, about 51,723 citation events across roughly 8,917 unique domains. Top-10 concentration by vertical came out at healthcare 71.2%, legal 54.7%, insurance 47.9%, fintech 45.3% and SaaS 28.4%.

Read that spread carefully, because it dictates strategy. Healthcare behaves like an oligopoly — NIH (14.7%), Mayo Clinic (12.3%) and the CDC (9.8%) lead, and a challenger brand's realistic play is being referenced within those sources rather than displacing them. B2B SaaS behaves like a long tail: G2 (9.4%), Reddit (6.8%), Capterra (5.1%), HubSpot (3.9%) and TrustRadius (3.4%) lead, and no single domain is worth a quarter of anyone's budget. Fintech sits in between, led by NerdWallet (11.2%), Investopedia (9.7%) and Bankrate (6.4%).

Notice what's missing from every one of those lists: the 5W top 15. If you sell B2B software, a Reuters placement and a Wikipedia article are not the shortest route to being cited — a G2 profile with real review volume is. Attrifast's source-type mix across all 51,723 citations tells the same story from another angle: editorial reviews 22.3%, long-tail blogs and newsletters 19.3%, vendor sites 14.7%, Reddit 11.4%, news and press 9.1%, Wikipedia 8.9%, forums and Q&A 7.8%, academic and government 6.5%.

Prestige and citation frequency have also come apart. 5W's own May 2026 audit found that WSJ, The New York Times, Bloomberg and the Financial Times do not appear in ChatGPT's US top 20, while its Trade Press AI Index found tech answers led by PCMag, TechRadar and CIO.com, travel by Skift, healthcare by STAT and Endpoints News. As Torossian put it: "The most prestigious outlet and the most cited outlet are now two different publications." For a comms lead whose media list was built on masthead prestige, that's a reordering exercise, not a tweak.

What actually moves the needle, ranked by measured effect

Once you know which domains your category's engines cite, the next question is which kind of presence on them produces citations. Here the measured effect sizes vary by more than an order of magnitude.

One large, underappreciated factor is backlink strength — our analysis Do Backlinks Matter for AI Citations? The 2026 Data examines how backlinks correlate with citation rates and the measured effect sizes.

Review profiles are the largest single measured swing in the public literature. Seer Interactive, commissioned by Trustpilot, ran 804,491 AI responses covering 1,926 brands and 15,783 prompts across ChatGPT, Google AI Mode, Gemini and Perplexity, primarily in March 2026. Median AI citation rate by profile tier: no claimed profile 1%, minimal profile with 1–13 reviews 53.5%, active profile around 78.5%, optimized around 84.5%. Going from nothing to thirteen reviews is the cheapest visibility intervention in this entire post. Two caveats keep it honest: the study is vendor-commissioned, and it notes that factors like marketing spend and audience size may influence results even after cohorting by domain rating.

Third-party ranked lists are the volume play. Evertune's review of the 6,000 most-cited URLs per model across six engines for March and April 2026 found that 63% of roughly 400 million citations pointed to listicles, with 71–86% of those being ranked lists. Position inside them matters too, but less than presence: Peec AI's analysis of nearly 200,000 AI responses across eight engines from September 2025 to March 2026 found rank-1 placement lifted brand mention probability by 16.5 percentage points in B2B SaaS and 13.4pp in emerging MarTech, and pulled brands 1.17 positions earlier in B2B SaaS answers. Peec also found impact saturating "after just a few repeated placements" — the fifth strong placement is worth far less than the first. Their conclusion belongs in your planning doc: "Five strong placements in frequently AI-cited third-party sources will matter more than 50 placements in articles AI engines never retrieve." If that's the lever you want first, our guide to getting into the best-tools lists AI engines cite covers the qualification and pitching mechanics.

LinkedIn is a top-tier source, and it rewards people over brands. Semrush's study with LinkedIn, covering 325,000 unique prompts and 89,000 LinkedIn URLs from January–February 2026 data, found LinkedIn appearing in about 11% of AI responses on average — 14.3% on ChatGPT Search, 13.5% on Google AI Mode, 5.3% on Perplexity. On the first two, 59% of cited LinkedIn content came from individual creators rather than company pages. Articles supply 50–66% of citations, the cited band is 500–2,000 words for articles and 50–299 for feed posts, roughly 95% of cited content is original rather than reshared, and 54–64% is practical advice rather than promotion. Median engagement on cited posts is a modest 15–25 reactions — citation is not a popularity contest, which is unusually good news for a small team with one credible expert on staff.

Earned editorial is the bulk of the pool; paid is a rounding error. Muck Rack's May 2026 edition of What Is AI Reading?, analyzing 25 million-plus links across ChatGPT, Claude and Gemini responses in 17 industries, attributes 84% of AI citations to earned media and 27% to journalism, while paid and advertorial content sits at 0.3%. Press releases account for under 1% of citations even after growing fivefold since July 2025, and they cluster in industry-trend answers rather than "best of" queries. If wire distribution is a line item in your GEO budget, that's the number to size it against.

Four objections worth taking seriously

"Owned media is still 44% of citations." Fair, and the counter-evidence is real: Yext analyzed 6.8 million citations from over 1.6 million AI responses across Gemini, OpenAI and Perplexity and classified 44% as brand-owned websites, 42% as controllable listings and directories, 8% as influenceable reviews and social, and 6% as uncontrollable. That looks like a flat contradiction of Muck Rack's 84%-earned figure, and the reconciliation matters more than picking a winner: the two cover different query mixes and classify differently — Yext counts directory listings as brand-managed, Muck Rack treats them as third-party surfaces. The synthesis is that your own site is table stakes and rarely the bottleneck, while the fastest available gains sit on surfaces you influence but don't own.

"Being cited isn't being recommended." This is the sharpest objection, and the data backs it. Lily Ray analyzed 100 B2B "best [category] software" queries in Google AI Overviews across April, May and June 2026 and found brands' own self-promotional listicles cited 323 times — with the brand cited but not recommended in 224 of those cases, or 69%. Your page got used as a source to recommend your competitors. This is exactly why citation rate and mention rate need separate targets and separate reporting lines; they can and do move in opposite directions. It's also why Geoptimizer scores them as distinct components — mention rate at 35% of the visibility score, citation rate at 25% — rather than collapsing them into one number.

"These sources are unstable." Correct, and the volatility is severe. Reddit's overall AI citation share fell from 2.02% to 1.01% between October 2025 and January 2026, roughly halving in four months, even as its exclusive citation rate — queries where Reddit was the only source — rose 31%. Earlier, Semrush's 13-week study of 230,000+ prompts saw Reddit's ChatGPT citation frequency collapse from about 60% of prompt responses to about 10% inside September 2025, with Wikipedia falling from ~55% to under 20% in the same window. Semrush's Sergei Rogulin attributed it to platform intent: "I believe the main reason for the drop is an attempt to avoid over-citing on certain websites, to be less biased toward them." Add Reddit's $60 million-a-year Google licensing deal entering renewal talks, and the conclusion writes itself: diversify across five to eight sources and never let one domain carry your category.

"Isn't this just SEO with extra steps?" Partly, and saying so plainly makes the rest of the argument more credible. The Trustpilot study found that 99.5% of Trustpilot citations arrive through organic search ranking rather than the engine seeking Trustpilot out. Digital Applied's analysis of 1,000 AI Overviews in April 2026 found domain authority the strongest page-level correlate of citation (+0.61), with schema markup associated with a 2.3× higher citation rate and content over 2,500 words 1.6×. Third-party placement works because those pages rank. A placement on a page that ranks for nothing buys you nothing — which is the single best filter for qualifying a target publication before you pitch it.

One more piece of context on where the pressure is highest: the same AI Overviews study found definitional queries averaging 5.6 citations and commercial queries only 3.1. The concentration bites hardest exactly where buyers are.

A prioritization framework for the next quarter

Everything above collapses into four moves, in order.

  1. Measure your own board before you pitch anyone. The industry top 15 is a starting hypothesis, not your list. Run your buyers' actual questions and record which domains the engines cite when answering them — that list will look far more like Attrifast's vertical tables than like the 5W index. Build the prompt set from real buyer questions across the funnel, not from your keyword list.
  2. Pick five to eight sources spanning at least two engine profiles. Engines draw from genuinely different pools: ChatGPT leans on Wikipedia, Gemini and Perplexity on Reddit, Claude on PubMed Central, Google AI Mode on fandom.com, and Grok's pool is a different universe entirely — Reddit 16.3%, YouTube 15.1%, Facebook 13.9%, with Wikipedia at only 3.4% across 1.9 million-plus US queries in June 2026. A portfolio that wins ChatGPT can leave you invisible on Grok, so read the engine-by-engine citation evidence before you allocate.
  3. Set separate targets for citation rate and mention rate. Given the 69% cited-but-not-recommended finding, a report that shows only "we appeared as a source" can hide a losing quarter. Track how often your category's prompts return your brand as an answer, alongside how often your placements are used as sources.
  4. Re-measure on a rolling window, not a spot check. Reddit halved in four months; Wikipedia lost roughly two thirds of its ChatGPT prompt-response frequency inside a single month. A one-off screenshot from the week after a placement lands tells you almost nothing. A rolling window with a confidence band tells you whether the placement stuck.

That last point is where tooling earns its keep. Geoptimizer runs your buyer-intent prompts live on ChatGPT, Gemini, Claude and Grok with web search enabled, reports mention rate and citation rate as separate components of a published 0–100 formula, and holds the headline score on a seven-day rolling window with a confidence band because AI answers are nondeterministic. For the competitor side of the picture, our method for running a share-of-answers competitive analysis shows how to trace a rival's lead to the sources winning it for them.

What to watch over the next six months

Three things could reshuffle the board before this post turns a year old. Reddit's Google licensing renegotiation is the biggest — it is the number-one or number-two source on most engines, so any change to its access terms moves everyone's numbers at once. Second, review-site consolidation: G2's acquisition of Capterra, Software Advice and GetApp closed on February 5, 2026, and Omniscient Digital modeled the combined properties at 3.68% of bottom-of-funnel citations, up from G2's own 2.09% — one commercial relationship now covering a materially larger slice of B2B buying answers. Third, citation density is rising: Attrifast measured roughly 28% growth year over year to May 2026, with Claude up 50%. More citations per answer means more room below the top 15, which cuts against the concentration thesis over time.

None of that changes the play for this quarter. It changes how often you check.

FAQ

What are the most cited AI citation sources in 2026? Across engines, Wikipedia and Reddit lead consistently — together they take 25.1% of ChatGPT's citations in Similarweb's US data for January–February 2026, and 25.6% in Ahrefs' July 2026 sample. Below the top two, membership varies sharply by engine: Google AI Mode's most-cited domain is fandom.com, Claude leans on PubMed Central, and Grok's pool is dominated by Reddit, YouTube and Facebook.

Is the "68% of AI answers come from 15 domains" statistic accurate? It's directionally sound but engine-averaged and unaudited. It comes from 5W Public Relations' index, which synthesizes six studies without publishing weighting methodology. Independently measured, ChatGPT's top 15 reach about 56.7% and Google AI Mode's about 38.2% (our sums of Similarweb's published shares). Concentration is real; the specific number should always carry an attribution.

Should I prioritize third-party placements over my own website? For most teams, yes — but not exclusively. Yext's classification found 44% of citations pointing to brand-owned sites, so your own pages still matter. The reason third-party work usually wins on effort-to-result is that a single well-ranked listicle or review profile can be cited across thousands of buyer prompts, and the measured effect sizes (1% to 53.5% median citation rate for claiming a review profile) beat most on-site changes.

How many third-party sources do I actually need? Peec AI's data shows impact saturating after a few repeated placements, and their guidance is that five strong placements in frequently cited sources beat 50 in sources engines never retrieve. Five to eight, spread across at least two different engine citation profiles, is a defensible target — enough for diversification against volatility like Reddit's, without spreading a small team thin.

Do press releases get cited by AI engines? Rarely, as a share of the whole. Muck Rack's May 2026 data puts press releases at under 1% of total citations despite fivefold growth since July 2025, and paid or advertorial content at 0.3%. They index best in industry-trend answers rather than commercial "best of" queries, so treat wire distribution as a supporting tactic rather than a citation strategy.

The takeaway

Concentration in AI citation sources is real, engine-specific, and only partly addressable. Of ChatGPT's 63.6% top-20 share, roughly 37 points sit on domains a content or PR team can genuinely work — and in your category, that board probably looks nothing like the industry list everyone is quoting. The teams that win the next four quarters won't be the ones with the longest target media list. They'll be the ones who measured which domains their buyers' questions actually pull from, placed five to eight times well, and checked whether it stuck.

Start with the measurement. Run your buyer-intent prompts across ChatGPT, Gemini, Claude and Grok, see which sources the engines are actually citing in your category, and build the pitch list from that — the free plan covers all four engines, so you can have your own citation board before you send a single email.

Keep reading

See it on your own domain.

Free visibility check across ChatGPT, Gemini, Claude, and Grok — about 30 seconds.

Run the free check