· Updated · 15 min read · Geoptimizer Team
Gemini vs AI Overviews vs AI Mode: What to Track
- generative-engine-optimization
- ai-visibility
- measurement
- google-ai-mode
- ai-overviews
- gemini
Google's three AI answer surfaces cite almost entirely different sources. When Ahrefs compared AI Overviews and AI Mode responses to the same queries, only 13.7% of citations overlapped — even though the two answers averaged 86% semantic similarity. So the honest answer to "which Google AI surface should I track?" is: track them separately, because a single blended "Google AI visibility" number averages three different retrieval stacks into a figure that cannot be acted on. This guide explains how each surface retrieves and cites, what evidence exists for how far apart they are, what you can actually measure on each one, and how to build a tracking plan that stops producing contradictory reports.
The urgency is new. Similarweb data reported by TechCrunch on 27 July 2026 shows AI Overviews now appearing in 43% of searches, up from 15% a year earlier, while AI Mode visits grew from 126 million in June 2025 to 279 million in May 2026. On the Alphabet Q2 2026 earnings call, Sundar Pichai said the Gemini app now has 950 million monthly active users and that AI Mode has "surpassed 1 billion monthly active users" since its global expansion. Three surfaces, all at enormous scale, all answering with links — and all sourcing those links differently.
Three surfaces, three retrieval stacks
Before the measurement question, the mechanics. These are not three skins on one system.
AI Overviews: a summary layered on ranked results
An AI Overview appears above traditional results on a classic search results page, on a subset of queries where Google's systems decide a synthesized answer helps. Its raw material is what Google Search already retrieved and ranked for that query, which is why its citations look comparatively familiar to anyone with a rank tracker.
The evidence supports that intuition. Semrush's study of 5,000 keywords found AI Overviews had roughly 86% domain and 67% URL overlap with Google's top 10 organic results. If you rank, you are in the candidate pool. Overviews are also comparatively concise in their sourcing: the same study measured an average of three unique linked domains per AI Overview.
AI Mode: query fan-out, then synthesis
AI Mode is a separate conversational surface. Google's own Search help documentation describes the retrieval method plainly: AI Mode works by "dividing your question into subtopics and searching for each one simultaneously" — the technique commonly called query fan-out. Google's guidance for generative AI features defines a fan-out as "a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results."
That single design decision explains most of what follows. One user question becomes many machine questions, each with its own retrieval, and the answer is assembled from a much wider pool. SE Ranking's analysis of 10,000 keywords found AI Mode responses carried 12.6 URLs on average — 122,617 links across 9,734 triggered responses. Semrush found a sources sidebar with roughly seven unique domains in 92% of AI Mode responses, and measured only ~35% URL overlap with the organic top 10, roughly half the figure for AI Overviews.
AI Mode also changed under the hood recently. At Google I/O on 19 May 2026, Google made Gemini 3.5 Flash "the new default model in AI Mode for everyone globally" and added the ability to ask a follow-up straight from an AI Overview and flow into an AI Mode conversation. The two surfaces are becoming more connected in the interface while remaining distinct in retrieval — a combination that makes conflated reporting even easier to fall into.
The Gemini app: generative first, grounded on demand
The Gemini app is not a search surface with a chat interface bolted on. It answers from the model, and reaches for Google Search when the question benefits from it. The consequence for citation tracking is direct: Google's Gemini Apps help page states that "not all responses include related links or sources", and that if the Sources button is absent, no links were provided for that response at all.
Where the app does ground an answer, the plumbing resembles what developers can see in the Gemini API's Grounding with Google Search tool: the model decides whether searching improves the answer, issues its own search queries, and returns citation annotations that tie specific spans of the answer to source URLs. Anyone tracking "Gemini" through an API is measuring this surface — a grounded assistant answer — not an AI Overview and not AI Mode.
The overlap data: four studies, one conclusion
Four independent analyses have measured how much AI Overviews and AI Mode agree on sources. They used different samples, different dates, and different matching rules, and they landed in different places — but every one of them found the majority of citations diverging.
| Study | Sample | Citation overlap found |
|---|---|---|
| Ahrefs (Sept 2025 US data) | 540,000 query pairs for citation analysis | 13.7% of citations overlap; 16.3% among top-3 citations |
| SE Ranking (June 2025) | 10,000 keywords | 10.7% exact URL overlap; 16% domain overlap |
| Victorious (Sept 2025) | 1,540 real-world queries | 30–35% of AI Overview URLs also appeared in AI Mode; 77% of unique domains showed up in only one surface |
| Semrush (July 2025) | 5,000 keywords, 150,000+ citations | AI Mode ~35% URL overlap with organic top 10 vs ~67% for AI Overviews |
The spread between 10.7% and 35% is itself instructive — it reflects how much measured overlap depends on sample, matching at URL versus domain level, and the date of collection. But no methodology produced a number that would justify treating one surface as a proxy for the other.
Victorious added a detail worth pinning to the wall: across its 1,540-query sample, no query produced identical citation lists between the two surfaces, and AI Mode drew on roughly nine domains per query against 7.6 for AI Overviews.
Same answer, different sources — and why that wrecks reports
The most useful finding in the Ahrefs data is not the 13.7%. It is that low citation overlap sits alongside 86% average semantic similarity, with the two surfaces sharing an identical opening sentence only 2.51% of the time and AI Mode responses running roughly four times longer.
Read that carefully. The two surfaces usually agree about the substance of the answer, and usually disagree about who gets credit for it. For a reader, that is fine. For a brand, it is the whole game — being the cited source is the outcome you can measure, link to, and earn traffic from. A blended Google score mixes a surface where you might hold a citation with one where you do not, and hands you an average that hides the only fact you could act on.
This is the same statistical trap that makes two GEO tools disagree about the same brand, which we covered in why AI visibility tools give you different scores: when the underlying samples differ, the resulting numbers differ for reasons that have nothing to do with your website.
There is a second layer: within a single surface, sources move between runs. SE Ranking ran the same 10,000-keyword set three times on the same day and found only 9.2% exact URL overlap between the three datasets (14.7% at domain level). Any surface-level comparison has to account for that noise floor before anyone declares a trend — a diagnostic we walk through in our guide to separating model-version effects from measurement noise.
What each surface can actually tell you
Retrieval differences are one half of the problem. Measurement access is the other, and it is unevenly distributed.
| Surface | First-party data | Third-party sampling | Practical metric |
|---|---|---|---|
| AI Overviews | Search Console generative AI performance report (impressions) | Mature: SERP scrapers see the rendered overview and its links | Impressions, citation presence per tracked query |
| AI Mode | Same Search Console report | Harder: no public API, conversational and multi-turn | Impressions; sampled citation rate on a fixed prompt set |
| Gemini app | None | API grounding as a close proxy | Mention rate, citation rate, prominence across repeated runs |
AI Overviews and AI Mode: one Search Console report
Google's generative AI performance report is the first-party source for both Search surfaces. Its help documentation names the covered features as AI Overviews and AI Mode, and defines an impression as "how many times links to your site were shown to a user in a generative AI feature." It also notes that Search Console excludes data from Search Labs experiments, which are still in active development.
Two limits matter for planning. First, the documentation presents these features together and does not document a way to split impressions by feature — so a rise in the line does not tell you which surface rose. Second, the report is built around impressions; clicks, CTR and query-level detail are not part of what the help page describes. That makes it excellent for direction and poor for diagnosis, which is exactly the gap a sampled citation-rate metric fills. (Where the clicks from these surfaces land in your analytics — filed under Organic Search, not an AI channel — is covered in AI Visibility vs AI Traffic: Connect Your Score to GA4.)
Geography adds a third wrinkle. AI Overviews and AI Mode launched in France on 22 July 2026, roughly sixteen months after nine other European markets, with a domain-level Search Console toggle documented two days earlier. If your impressions curve for a country starts at zero this month, that is a rollout, not a ranking event — and only surface-and-country-level reporting will show you the difference.
The Gemini app: no first-party report, so sample it
There is no Search Console equivalent for the Gemini assistant. The Search generative AI control covers AI Overviews, AI Mode and generative AI features in Discover, and Google's documentation is explicit that it "isn't used as a ranking or inclusion signal affecting other parts of Search" — but the Gemini app is not in that list, and its impressions do not appear in your Search Console property.
So the only practical approach for the assistant surface is sampling: run a fixed set of buyer-intent prompts repeatedly, with search grounding enabled, and record whether your brand is mentioned, whether your domain is cited, and how prominently. That is what Geoptimizer's published methodology describes for its Gemini score — Gemini queried with Google Search grounding, scored as mention rate, citation rate, prominence and sentiment over a 7-day rolling window — and it states plainly that API answers "approximate but do not exactly equal what users see in the consumer apps." That caveat is the correct one to carry into any report: a sampled assistant score is a good directional read on the Gemini surface, and it is not a measurement of AI Overviews or AI Mode.
If assistants mischaracterize your offering, how to fix ChatGPT describing your brand wrong walks through practical prompts, grounding tactics, and content fixes to correct brand descriptions at scale.
Five rules for a Google AI tracking plan that doesn't contradict itself
1. Never average across surfaces. One blended "Google AI visibility" figure combines an impressions-based Search metric with a sampled assistant metric. When it moves, no one can say which surface moved or why. Report three lines, or report one line and name the surface it belongs to.
2. Match the metric to what the surface exposes. Impressions for the Search surfaces, because that is what Google gives you. Citation rate across repeated runs for the assistant, because that is what sampling gives you. Trying to force both into the same unit invents precision that does not exist.
3. Fix the sampling window before you compare anything. With 9.2% run-to-run URL overlap inside a single surface, a snapshot is a snapshot. A rolling window with a confidence band is the minimum honest presentation, and any two numbers you compare should come from equal-length windows.
4. Segment by country and device. The France launch is the clean example, but every surface rolls out on its own schedule; AI Mode's help page notes support across the Americas, Asia-Pacific and EMEA in over 90 languages. Global averages will keep absorbing local step-changes and presenting them as trends.
5. Write the surface name into every reporting line. "Google AI visibility up 6 points" is unusable. "Gemini assistant citation rate up 6 points on 25 tracked prompts; Search generative AI impressions flat" is a sentence someone can act on.
Which surface deserves your attention first
Prioritization depends on where your buyers actually ask, and there is no universal answer. A workable ordering:
Start with AI Overviews if your demand is query-shaped. For head and mid-tail informational queries where people search rather than converse, AI Overviews now appear on 43% of searches per the Similarweb data above, and their citation pool overlaps heavily with organic rankings. The work is largely continuous with SEO, and the payoff shows up in the Search Console impressions line.
Prioritize AI Mode if your category involves comparison and research. Fan-out rewards depth across subtopics rather than one strong page for one head term. If your buyers ask multi-part questions — "which tool handles X for a team of Y under Z budget" — AI Mode's twelve-plus source slots per answer are a wider door than an AI Overview's three domains, and a different set of pages will fit through it.
Prioritize the Gemini app if your buyers work inside assistants. At 950 million monthly active users, the Gemini app is a genuine discovery surface, particularly for research and drafting workflows where a recommendation is requested rather than a link. It is also the surface with no first-party reporting, so if you do not sample it, you have no data on it at all.
For most teams the honest allocation is: all three, weighted, with the Search surfaces tracked through Search Console and the assistant tracked through repeated prompt sampling. The mistake is not choosing wrong — it is choosing one and reporting it as though it covered Google.
For teams that want a concrete timeline and tasks to operationalize this mix, use the 90-day plan for generative AI optimization in B2B SaaS to map out what to test, measure, and scale.
What moves all three
The surfaces diverge on retrieval, but they share a foundation, and Google says so directly. Its AI features optimization guide states that "the best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems." The same page is equally direct about a popular shortcut: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search," and such files are ignored by Google Search. It also cautions that generating content specifically to target fan-out query variations can fall under its scaled content abuse policy.
Three things travel well across all three surfaces:
- Retrievability. Being in the candidate pool is upstream of everything. The ~86% domain overlap between AI Overviews and the organic top 10 means classic ranking work still buys you the ticket to the AI Overview draw.
- Extractable passages. Fan-out retrieval matches sub-questions, not page titles. Content organized so that a single self-contained paragraph answers a single specific question gives all three systems something to lift and attribute.
- Coverage of the whole question space. AI Mode's dozen sources per answer come from a dozen sub-queries. Pages that thoroughly cover adjacent subtopics — pricing, limits, comparisons, edge cases — are eligible for more of those slots than one broad overview page.
And one thing worth remembering when you present all of this to a stakeholder: appearing in an AI answer is not the same as receiving a click. Pew Research's panel study of 900 US adults found that in March 2025, users clicked a traditional result in 8% of visits where an AI summary appeared, versus 15% without one, and clicked a link inside the summary in just 1% of visits. Citation share is a visibility metric first and a traffic metric second.
FAQ
Is AI Mode the same as AI Overviews? No. Google's help documentation describes AI Mode as expanding "what AI Overviews can do with more advanced reasoning," and it retrieves by splitting your question into subtopics and searching each simultaneously. AI Overviews summarize results Google already ranked for the single query. Independent studies put their citation overlap between roughly 11% and 35%, depending on methodology.
Does tracking Gemini tell me if I appear in AI Overviews? Not reliably. Tools that track "Gemini" typically query the assistant with Google Search grounding, which is a different surface from an AI Overview on a search results page. Use Search Console's generative AI performance report for the Search surfaces and sampled prompt runs for the assistant.
Can Search Console separate AI Mode impressions from AI Overviews? Google's help page for the generative AI performance report names both features as covered and does not document a way to split impressions by individual feature. Treat the report as one combined Search-side signal, and use country, device and page dimensions to interpret movement.
Why do my AI visibility numbers change between scans even when nothing changed on my site? Because these systems sample rather than rank. SE Ranking re-ran the same 10,000 keywords three times in one day and found only 9.2% exact URL overlap between the runs. Rolling windows and confidence bands exist precisely to make that noise legible instead of alarming.
Should I optimize differently for each surface? The foundation is shared — Google states its generative AI features are rooted in core Search ranking systems. The differences are emphasis: AI Overviews reward strong ranking on the head query, AI Mode rewards breadth across the sub-questions a fan-out generates, and assistant surfaces reward being described accurately across the wider web that grounding retrieves.
Track three surfaces, report three numbers
Conflicting Google AI reports are rarely a tooling failure. They are what happens when three retrieval stacks — a ranked-results summary, a fan-out synthesis, and a grounded assistant — get averaged into one figure. Split them, match each metric to what the surface actually exposes, fix your window, and the contradictions mostly disappear.
If you want a baseline for the assistant side of that picture today, run a free AI Visibility Check — it asks Gemini, ChatGPT, Claude and Grok live, with web search on, whether they mention and cite your brand, and returns a scored snapshot in about 30 seconds with no signup. Pair that with your Search Console generative AI report, and you will have the two halves of Google covered separately, which is the only way either half means anything. When you are ready to track it continuously across a real prompt set, the plan comparison shows what each tier includes.