· Updated · 18 min read · Geoptimizer Team
Get Cited by AI: The 2026 Evidence, Engine by Engine
- generative engine optimization
- ai citations
- chatgpt
- gemini
- claude
- grok
On May 7, 2026, ChatGPT started putting prominent, clickable brand links inside its answers. In the weeks that followed, Similarweb's desktop clickstream panel measured total ChatGPT referrals up 157.7% week-over-week, with referrals landing on brand homepages up 354.7% (Similarweb). Being named in an AI answer stopped being a vanity metric that afternoon.
"Get cited by AI" is now three separate outcomes with three different mechanics — being mentioned, being cited, and being clicked — and the four big engines reach for such different source pools that one generic playbook can be measurably wrong for at least one of them. Here's what the 2026 data supports, engine by engine, including the popular tactics that show no effect in controlled tests.
TL;DR:
- Mentioned, cited, and clicked are decoupled outcomes — and engines barely share sources (Claude and ChatGPT overlap on just 13% of cited domains).
- Earned media dominates: roughly 84% of cited links are pages you don't own, so off-site presence beats on-page tinkering.
- The one strong on-page lever is answering the exact buyer question in the top third of the page, in cleanly liftable prose.
- Pricing pages are the top-cited page format; self-published listicles score negative.
- Schema markup, llms.txt, and Core Web Vitals show no measurable citation effect in matched-control tests — fix crawler access instead, and remember no major AI crawler executes JavaScript.
Mentioned, cited, clicked: three different outcomes
A mention is the engine naming your brand in its prose. A citation is the engine linking a page as a source. A click is a human leaving the answer for your site. These decouple more than people expect.
Semrush's 2026 AI Visibility Index — built on 126 million US prompts from January to April 2026 across 1,200+ brands — measured how often the brands named in an answer were also the brands cited as sources. That overlap ran from 64% on Google AI Overviews down to 30% on Gemini. As the write-up puts it, a brand can be "mentioned constantly and cited almost nowhere, or cited constantly while almost never being the brand the AI is actually discussing" (PPC Land's analysis of the index).
Before May 2026, a mention without a citation was mostly a branding win. Now it moves traffic. Profound's read on the same OpenAI change found referrals to monitored brand sites up 60–65% overnight, the homepage share of ChatGPT referrals jumping from 4% to 24%, and the share of ChatGPT responses containing URLs rising from about 4.5% to 20–24%. B2B software and SaaS saw daily referrals rise more than 200%; e-commerce and retail were essentially flat (Profound). Similarweb's own summary of the pattern: "Rather than eliminating clicks, AI may simply be redirecting them toward brand homepages, discovered through conversational queries."
This is exactly why a single blended "AI visibility" number is hard to act on. If mentions are strong and citations are weak, your fix is content and technical access on your own domain. If citations are strong and mentions are weak, engines are reading you as a source but not as an answer to the buying question. Those are different projects. Geoptimizer splits the score into its parts for that reason — mention rate 35%, citation rate 25%, prominence 20%, sentiment 20%, calculated per engine and then averaged, with the full formula and version history published on the methodology page rather than kept behind a dashboard.
For a practical, evidence‑backed way to reconcile those tradeoffs and translate per‑engine signals into a single, actionable metric, see the evidence behind our AI Visibility Score formula.
The four engines do not read the same web
The single most useful fact for 2026 planning: cross-engine source pools barely overlap.
Otterly.ai analysed 379,321 Claude citations across 16,406 domains in June 2026 on a SaaS and technology prompt set. Claude and ChatGPT shared only 13% of cited domains and 4.2% of cited URLs (Otterly.ai). Semrush's index found citation sets changing roughly 50% month-over-month, with about 11% cross-platform overlap.
The Reddit contrast makes it concrete. Reddit is the #1 domain in Gemini at 29.2% mention share across 3M+ US queries, and #1 in Grok at 16.3% across 1.9M US queries (Ahrefs Brand Radar, June 2026: Gemini, Grok). In Otterly's 379,321-citation Claude dataset, Reddit appeared exactly zero times.
Here is the cheat sheet, with scope attached to every number:
| ChatGPT | Gemini | Claude | Grok | |
|---|---|---|---|---|
| Responses containing citations (Muck Rack, 25M+ links) | 96% | 82% | 55% | not measured |
| Avg citations per response | 5 (Muck Rack) / 15.4 (Semrush) | 8 (Muck Rack) / 3.3 (Semrush) | 13 (Muck Rack) | not measured |
| Top cited domain | Wikipedia | Reddit (29.2%) | PubMed Central | Reddit (16.3%) |
| Brand-owned URL share (Discovered Labs, 2M citations) | 39% | 14% | — | — |
| Median age of cited page | 8.0 months | 7.8 months | 5.1 months | — |
Two caveats. Semrush and Muck Rack disagree on ChatGPT's citation density (15.4 versus 5) because they use different prompt sets and different definitions of a citation — surfaced link versus underlying retrieved source. Both are large datasets; don't average them, note which question each answers. And most of these studies come from companies selling adjacent tooling, which doesn't make the numbers wrong but does make the methodology section worth reading.
A few more numbers that change what you'd do:
- Claude gives 64.0% of its citations to brand, product and company domains, with news and media at 14.9% and social media at 0.9% (Otterly, SaaS/tech prompts). Its long tail is enormous: the top 10 domains account for just 9.5% of citations, and 32.9% of cited domains were cited exactly once.
- Gemini's pool is the most concentrated. Reddit, YouTube and Wikipedia together take more than 55% of mention share, and Gemini is the least likely to cite brand-owned pages (14%).
- Grok is the smallest of the four by web traffic share — 2.8% of worldwide generative-AI website visits in May 2026 versus ChatGPT's 52.7%, Gemini's 27.3% and Claude's 8.9% (Similarweb data via PPC Land; websites only, excluding apps) — but Ahrefs cites SpaceX's May 2026 SEC filing reporting 117 million people using Grok's features monthly. Its citation pool is user-generated content generally: Reddit 16.3%, YouTube 15.1%, Facebook 13.9%, Instagram 5.9%, Quora 5.5%, TikTok 4.8%. Notably, x.com itself was only #12 at 1.4% mention share, though it climbed 15 places month-over-month — and Grok separately retrieves live X posts through xAI's X Search tool, a surface web-citation trackers may not fully capture.
With 13% domain overlap between the two engines that were compared directly, checking one engine tells you very little about the other three. That's the practical case for tracking all four rather than sampling one.
What the evidence says actually moves citations
Two independent lines of research — one correlational, one a controlled multivariate model — land in the same place: off-site presence outweighs on-site tinkering.
Ahrefs' study of AI Overview brand mentions across 75,000 brands found branded web mentions correlating at 0.664, well ahead of branded anchors (0.527), branded search volume (0.392), Domain Rating (0.326) and backlinks (0.218) (Ahrefs). Ahrefs' Ryan Law summarised the implication: "Unlinked mentions — text written about your brand on other websites — have very little impact on SEO, but a much bigger impact on GEO… LLMs derive their understanding of a brand's authority from words on the page."
Discovered Labs' regression across 2 million citations from four engines (six months to April 2026) reached the same conclusion from the other direction: AI-perceived domain authority was roughly 6× as influential as the strongest individual page-level feature — mean absolute SHAP 0.38 versus 0.06 (Discovered Labs). Their conclusion for challenger brands is that off-page authority sits upstream of every on-page lever.
Where does off-page presence come from? Mostly earned media. Muck Rack's Generative Pulse analysed 25M+ links from ChatGPT, Claude and Gemini across 17 industries and found 84% of cited links were earned media in the May 2026 edition (82–89% across three editions since July 2025), with journalism alone at 27% and paid or advertorial content at 0.3% (Muck Rack). Query type matters: industry-trend questions drive journalism citations at twice the rate of how-to questions, and press releases appear 3.5× more in trend responses than in "best of" queries.
The Semrush index shows what that looks like for a specific brand. Patagonia held an AI visibility score of 79–80 across the study with roughly 87,000 mentions — and the dominant driver was third-party outdoor-gear review sites. Reddit, OutdoorGearLab and REI generated 65,000+ mentions between them, more than all Tier-1 traditional media combined. Shopify built the deepest citation footprint measured (42,100 cited pages) on three layers: specialist review platforms like G2 and Capterra, its own documentation (222 doc topics scoring above 90 on the visibility scale), and community sources including 76,100 YouTube citations and 44,500 from Reddit.
The one on-page lever that survives controls
It isn't a technical setting. In the Discovered Labs model, prompt–content alignment carried a standardised β of +0.37 (95% CI +0.33 to +0.41) — about 3× larger than the next-strongest page-level signal, with a one-standard-deviation improvement associated with roughly 30% more citations. Supporting effects: page length +0.13, title–prompt similarity +0.09, an FAQ section +0.07 (domain-conditional), a TL;DR block +0.05.
Two details make that actionable. The median depth of the paragraph engines actually cite is 0.36 on a 0-to-1 top-to-bottom scale — engines pull from the top third of the page. And the page-format coefficients are genuinely counterintuitive: pricing pages score +0.39, the largest positive of any format, while listicles and reviews score −0.12. Comparison and how-to pages sit near the reference category.
So the brief is: answer the exact question a buyer asks, near the top of the page, in prose that can be lifted out cleanly — and stop treating your pricing and comparison pages as bottom-of-funnel-only assets. That's a content decision, not a plugin.
What the checklists get wrong
Three staples of the standard GEO checklist have now been tested, and the results are worth knowing before you spend a sprint on them.
Schema markup. Ahrefs ran a matched-control study on 1,885 pages that added JSON-LD between August 2025 and March 2026, against 4,000 control pages. The measured effect on citations: Google AI Overviews −4.6% (a small but statistically significant decline), AI Mode +2.4% and ChatGPT +2.2% (both indistinguishable from zero) (Ahrefs). The widely quoted correlation is real — across 6 million URLs, AI-cited pages were almost 3× more likely to carry JSON-LD — but schema lives on better-maintained sites that do many other things well. As the study puts it: "If a page is already getting picked up, our data suggests that adding schema isn't going to push it higher."
llms.txt. SE Ranking looked at roughly 300,000 domains and found no measurable relationship between having an llms.txt file and citation frequency; removing the feature from their predictive model actually improved its accuracy (SE Ranking, covered independently by Search Engine Journal).
Core Web Vitals. In the Discovered Labs model, real-user LCP, INP and CLS, synthetic Lighthouse scores, and outbound link count all showed no significant independent effect on citations.
Google's own documentation says much the same thing in plainer language: "You don't need to create new machine readable files, AI text files, markup, or Markdown," there's "no requirement to break your content into tiny pieces for AI," and "Structured data isn't required for generative AI search" (Google Search Central). Google's AI features page adds that there are "no additional technical requirements" beyond being indexed and eligible for a snippet (Google).
None of these tactics are scams. Schema still powers rich results, llms.txt costs twenty minutes, and fast pages are good for humans. They're just unproven for this specific job, and they're a poor substitute for the two things that do show effects: off-page presence and answer-shaped content.
Technical access is the real exception — and it gets harder in September
There is one technical category with unambiguous evidence behind it: whether the engines can read your pages at all.
No major AI crawler executes JavaScript. Vercel's crawler analysis found ChatGPT and Claude bots fetch JS files (11.50% and 23.84% of requests respectively) but never run them, so client-side-rendered content is effectively invisible (Vercel). That study is dated December 2024 and the crawler fleet has changed since, so treat it as the best available evidence and verify your own rendered HTML rather than assuming.
Blocking the wrong bot costs you a specific engine. OpenAI runs several agents with separate purposes, and its documentation is explicit that sites disallowing OAI-SearchBot "will not be shown in ChatGPT search answers, though can still appear as navigational links" (OpenAI). GPTBot (training) and ChatGPT-User (user-triggered fetches) are separate controls. Anthropic likewise runs three: ClaudeBot for training, Claude-User for fetches triggered by a user's question, and Claude-SearchBot, which "navigates the web to improve search result quality for users" (Anthropic). A blanket disallow written in 2024 to keep your content out of training data may be quietly removing you from answers today.
That risk gets bigger soon. From September 15, 2026, Cloudflare will block "mixed-use" AI crawlers by default on pages that host ads — applying to new customers, new sites of existing customers, and all existing free-tier customers (TechCrunch). Plenty of site owners will change their AI visibility without ever making a decision about it. We've written a bot-by-bot decision framework for that deadline in Cloudflare's Sept 15 AI Crawler Rules: Who to Let In — worth reading before the defaults land.
If you want the five-minute version, the free AI Crawler Checker reads your robots.txt and llms.txt and tells you which of the 10 AI crawlers that matter — GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, Google-Extended, GrokBot, PerplexityBot and CCBot — can currently reach your site. No signup, results in seconds.
For a short, practical walkthrough that shows how to create an llms.txt and grant AI crawlers access in about twenty minutes, see llms.txt and AI Crawler Access: The 20-Minute Setup.
A per-engine playbook
Generic advice loses on at least one engine. Here's what the source data implies for each.
ChatGPT
Highest citation rate of the four (96% of responses per Muck Rack) and the most willing to cite brand-owned pages: 39% of citations raw, 53% position-weighted, per Discovered Labs. It's also the most tolerant of older content, with a median cited page age of 8.0 months. Priorities: keep your own high-intent pages crawlable and answer-shaped, get into Wikipedia-adjacent reference material and news wires honestly, and check OAI-SearchBot specifically rather than assuming GPTBot covers it. Since May 7, a mention now converts into a homepage visit — so brand-name consistency across the web is a traffic lever, not just a positioning exercise.
Gemini
Cites the fewest sources per response (3.3 in the Semrush index) but names among the most brands (4.7 average). Its pool is dominated by Reddit, YouTube and Wikipedia, and it's the least likely to link your own site. Google's published guidance amounts to "keep doing good SEO" — Search Central states that best practices for SEO "continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems." For retail and local, Merchant Center feeds and Business Profiles are the Google-specific extras that do matter. One honest disagreement in the data: Ahrefs' all-topic query set puts Reddit at 29.2% mention share for Gemini, while Tinuiti's nine-category consumer tracking reports social citations are far rarer on Gemini than on AI Overviews (Tinuiti). Different methods, different definitions — check your own category rather than trusting either number universally.
Claude
Cites in only 55% of responses, but cites deeply when it does (13 sources on average). It overwhelmingly favours brand, product and company domains (64.0%) and high-authority institutional sources, and it has the strongest freshness preference — median cited page 5.1 months, with 60% of citations under six months old. Your documentation, product pages and technical explainers are the Claude play, and they need updating on a schedule. Social barely registers, and Reddit didn't appear at all in Otterly's SaaS/tech dataset.
Grok
The only engine with a retrieval surface built on live social posts. Its web citations skew heavily to user-generated content — Reddit, YouTube, Facebook, Instagram, Quora, TikTok — so breadth across communities matters more than depth on any one. Because Grok additionally pulls X posts via X Search with handle-level filters, an active, entity-consistent X presence (consistent brand name, consistent description of what you do) is worth maintaining even though x.com ranks only #12 in its web-citation profile today.
Across all four, for B2B
Profound's analysis of 1.4 million citations across six AI surfaces found LinkedIn is the #1 most-cited domain for professional, B2B and software queries, moving from roughly #11 to #5 on ChatGPT between November 2025 and February 2026. The content-type mix shifted too: profiles fell from 33.9% to 14.5% of LinkedIn citations while posts rose to 26.0% and long-form articles to 8.9% (Profound). The practical read: published posts and articles, not a polished company page.
Where the honest limits are
Four things a credible GEO plan should account for.
The founding "40% visibility boost" number is weaker than its reputation. A July 2026 scoping review of 45 studies lists five qualifications: the gain was measured with the source already placed in a fixed five-document context, the attention metric assumed earlier citations get more attention without user validation, the LLM judge came from the same model family as the generator, there was no downstream validation of clicks or conversions, and it captured a single point in time. The review's verdict is that the original paper's contribution was naming the phenomenon, "not providing a generalizable recipe" (Martinez, 2026; original paper).
Optimisation can backfire upstream. In the same review, an end-to-end test found body-only optimisation reduced average top-20 retrieval presence by about 9%, top-10 presence by 16% and final citation by 6%. In another benchmark, only 3 of 54 method–domain combinations showed significant positive effects. A rewrite that wins inside a fixed context can lose you the retrieval step that gets you into the context at all.
Single measurements are noise. Daily source-level Jaccard similarity across four engines ran 0.34–0.42 over 45 days, and the review recommends 7–8 repetitions per query as a starting point for stable measurement. That's the arithmetic behind reporting a rolling window with a confidence band instead of a single number, and labelling one-off runs as snapshots.
Manufactured mentions are a losing bet. MediaPost reported in July 2026 on companies flooding Reddit with AI-generated posts to manipulate citations. Reddit's countermeasures now block roughly 23 million spam views and catch about 25,000 spammy posts and comments daily, revoking nearly 2 million inauthentic votes per day, with spam exposure down 20% between January and March 2026 (MediaPost). Google's AI optimisation guide lists pursuing inauthentic mentions among the things that aren't helpful. The durable version of this work is being genuinely worth mentioning.
One more piece of context: Semrush surveyed 481 marketers and found 45% cannot accurately measure brand visibility in AI answers, and only 9% have tools tracking all the relevant metrics. Teams with integrated SEO and AI workflows reported increased traffic or leads at 81%, versus 36% for siloed teams — the largest single factor in that research.
FAQ
Is being mentioned better than being cited? They serve different purposes and you want both. A citation is a link an engine attributes as a source; a mention is the engine naming you in its answer. Since ChatGPT's May 2026 change, mentions can send homepage traffic on their own, which is why mention rate and citation rate belong in a score as separate components rather than blended into one figure.
Do I need an llms.txt file to get cited by AI? The evidence says it's optional. Across ~300,000 domains, SE Ranking found no measurable relationship between llms.txt and citation frequency, and Google states you don't need AI text files at all. It's cheap to add and harmless, but crawler access in robots.txt is the setting with documented consequences — blocking OAI-SearchBot, for example, removes you from ChatGPT search answers.
How long does it take to see a change? Different levers move on different clocks. Crawler access changes take effect as soon as bots re-fetch. Content changes depend on re-crawl and re-retrieval. Off-page authority — the strongest factor in the data — moves on the timeline of earned media and community presence, which is months. Given daily Jaccard instability of 0.34–0.42, judge any change over a multi-week window, not a single scan.
Should I check all four engines, or is ChatGPT enough? Check all four. Claude and ChatGPT shared only 13% of cited domains in Otterly's June 2026 dataset, Reddit is the top domain in Gemini and Grok but absent from that Claude dataset, and citation sets turn over roughly 50% month-over-month. One engine tells you very little about the other three.
Start with what you can verify today
The short version of the 2026 evidence: earn mentions off your own site, answer the exact buyer question near the top of your own pages, keep every relevant crawler unblocked, and measure per engine over a window rather than per scan. The tactics that look most like a checklist — schema, llms.txt, performance scores — are the ones the controlled tests can't distinguish from zero. And if the engines are not just omitting you but describing you incorrectly, that is a different repair job — the diagnosis and fixes are in ChatGPT Describes Your Brand Wrong: How to Fix It.
You can see where you stand in about 30 seconds. The free AI Visibility Check runs buyer-intent prompts live on ChatGPT, Gemini, Claude and Grok with web search enabled and returns a snapshot score showing which of them mention and cite you — no signup, no credit card. If you want continuous tracking, the free plan covers 3 AI-suggested prompts on all four engines forever, and paid plans start at $19/month with every engine included at every tier.