· Updated · 19 min read · Geoptimizer Team

AI Visibility vs AI Traffic: Connect Your Score to GA4

  • ai-visibility
  • ga4
  • attribution
  • measurement
  • generative-engine-optimization
AI Visibility vs AI Traffic: Connect Your Score to GA4

An AI visibility score and a GA4 report answer two different questions on purpose: the score measures how often engines say your name inside an answer, while GA4 measures the narrow subset of that exposure that produced a click it could label. The gap between them is enormous and well documented — Ahrefs found AI search was 0.5% of its visits but 12.1% of its signups over a 30-day window, and Similarweb found that when ChatGPT recommends a brand, roughly 56% of the resulting visits arrive via branded search rather than as an AI referral at all. So if your GA4 says AI sent you nothing, the honest reading is not that GEO isn't working. It's that GA4 is structurally incapable of showing you most of the return, and you have to instrument around it.

The two questions, and why one number can't answer both

A visibility score is a measure of exposure inside the answer. It asks: across the prompts your buyers actually type, how often does an engine mention you, cite you, place you prominently, and describe you favorably? That question has an answer whether or not anyone clicks.

GA4 asks something narrower. A session only enters its report if the user clicked, the browser passed something GA4 could classify, and the click landed inside a window GA4 can attribute.

Three structural facts make those two numbers permanently different:

  1. Most AI-influenced demand doesn't arrive as an AI click. Similarweb's downstream study found users who received a ChatGPT brand recommendation were "2.5 times more likely to visit that brand's website in the seven days that followed," with search accounting for closer to 56% of those visits versus around 40% for comparable visits. In GA4 that click is Organic Search on a branded query: "When a user asks ChatGPT for a product recommendation, reads the answer, and then searches for that brand three days later, your analytics has no idea those two things are connected."
  2. Google's own AI surfaces are filed as Organic Search by design. More on this below — it's a definition, not a bug you can configure around.
  3. A meaningful share of real AI clicks never gets labeled correctly, because of a UTM precedence rule most channel setups get wrong.

You are not alone in the gap. Semrush's 2026 AI Visibility Index, built on 126 million US AI search prompts from January–April 2026, reports that 45% of marketing leaders cannot accurately measure their brand visibility in AI-generated answers and only 9% have tools covering all relevant metrics. Conductor's State of AEO/GEO survey of 250+ US respondents at 500+ employee organizations found 94% planning to increase AEO investment in 2026, averaging 12% of digital budgets, with "difficulty measuring ROI of AEO efforts" a top-three challenge. The budget exists. The measurement is what's missing.

What GA4 measures natively — and exactly where it stops

On 13 May 2026 GA4 added a native AI Assistants default channel, with broad rollout through early June. Google's channel group documentation defines it verbatim:

"AI Assistants is the channel by which users arrive at your site from sources like ChatGPT, Gemini, Deepseek, Copilot, or Grok. It excludes Google's AI Overviews and AI Mode."

Four limits are baked into that sentence and the rules behind it.

The named-source list is incomplete. Claude and Perplexity are not on it, so traffic from either never reaches the native channel.

Google's own AI surfaces are explicitly elsewhere. The same documentation defines Organic Search as the channel by which users arrive "via non-ad links in organic-search results, including Google's AI Overviews and AI Mode." No GA4 configuration isolates those clicks. That is the largest measurement hole in the stack, and it is a design decision. (What each of Google's three AI surfaces exposes instead, and which metric fits each, is broken down in Gemini vs AI Overviews vs AI Mode: What to Track.)

The matching rule keys on medium. The rule is that the medium exactly matches ai-assistant, set when the referrer matches Google's internal list — which brings us to the trap in the next section.

Classification is forward-only. GA4 stamps the medium at collection time, so nothing that arrived before 13 May 2026 gets reclassified. An "AI channel growth curve" starting on 13 May is a rollout date, not a trend. Custom channel groups behave differently — Google's channel group docs confirm they can be applied retroactively — which is how you build a trend line long enough to matter. Standard properties get two custom groups (360 gets five), and they can't be used in the Key events paths report.

The 35–70% GA4 hides, sourced honestly

You will see "GA4 hides 35–70% of your AI traffic" repeated a lot. It's directionally right, but a CFO who checks will find the range stacks two different measurements.

The low end. Clickport's April 2026 analysis covered 371,847 sessions, of which 2,619 were AI sessions. Of those, 1,684 arrived with a referrer and "the other 935 came in UTM-only with no referrer" — 35.7%, landing in Direct or Unassigned by default. The platform split of that AI traffic: ChatGPT 1,994 sessions (76%), Perplexity 467 (18%), Claude 78 (3%), Gemini 42 (2%), Copilot 31 (1%).

The high end. Loamly's February 2026 update reports that "14,413 out of 20,428 total AI visits arrive without referrer headers. That is 70.6%, not 60%" from a base of 446,405 visits. Read the method before you quote it: detection uses a five-layer stack whose cryptographic-verification layer carries 0.99 confidence but whose navigation-timing paste-detection layer runs at 0.50–0.65.

The defensible phrasing: published measurements of referrer-less AI sessions cluster between roughly a third and two thirds, depending on dataset and detection method. Those are not the same denominator, and neither is literally "the percentage GA4 hides."

The UTM precedence trap

Here is the mechanic that matters more than the range. If a URL carries any UTM parameter, GA4 classifies on the UTM values and ignores the referrer. ChatGPT tags its links utm_source=chatgpt.com with no utm_medium. So GA4 has a source, has no medium, matches no channel rule, and files the session under Unassigned — not Referral, not AI Assistants. Google's documentation defines Unassigned as "the value Analytics uses when there are no other channel rules that match the event data."

The consequence is that a single source can fragment across three channels in the same week — chatgpt.com / ai-assistant, chatgpt.com / referral, and chatgpt.com / (not set) — because GA4 decides the channel using source and medium together, and the medium disappears on app and in-app-browser sessions.

The good news, which has aged the "hidden traffic" framing: those referrer-less ChatGPT clicks are still identifiable by source. They're hiding in Unassigned, not truly invisible. Genuinely unrecoverable are (a) AI Overviews and AI Mode clicks, by Google's own definition, (b) referrer-less clicks from assistants that don't tag, and (c) copy-paste and searched-the-brand-later journeys. Post-May-2026, the biggest hole is Google's surfaces, not ChatGPT.

Server-side tracking is oversold as the remedy. It recovers referrers lost client-side — redirect chains, JS timing, consent tooling — but not a referrer the browser never sent, and rel="noreferrer" links and native-app webviews send no header to any server. Adopt Seer Interactive's framing instead: any AI traffic number is "a directional floor rather than a precise count."

The six-layer build: from a broken channel report to a signup number

This is roughly a day of work, most of it in the first thirty minutes.

Layer 1 — A retroactive custom channel group

Admin → Data display → Channel groups → create a custom channel group. Build your AI channel on Session source matches regex, and place the rule above Referral in the ordering — if Referral sits above it, AI sessions get claimed by Referral first. Add a second condition so the native channel folds in: Default channel group exactly matches AI Assistants. Because custom groups apply retroactively, you get AI-channel history from before 13 May 2026 — history the native channel will never produce.

Layer 2 — A boundary-aware regex on source, never medium

Search Engine Journal publishes a pattern that won't match a domain that merely contains the token:

.*(^|[/.:@?&=])(chatgpt\.com|chat-gpt\.org|openai\.com|perplexity|gemini\.google\.com|copilot\.microsoft\.com|edgepilot|edgeservices|claude\.ai|deepseek\.com|grok\.com|you\.com|nimble\.ai|iask\.ai|aitastic\.app|bnngpt\.com|writesonic\.com|copy\.ai)([/.:@?&#=]|$).*

Their warning is worth repeating: "Never throw in a bare token like gpt on its own, because it'll match any source that happens to contain those three letters and drag in false positives." Note also that GA4 regex matching is case-sensitive, so include the case variants you actually see in your data.

Because the rule keys on source, it rescues the UTM-only ChatGPT sessions sitting in Unassigned — which is the bulk of the recoverable volume. Re-review the pattern quarterly: assistants add domains constantly, and Google's own named-source list has already changed once.

Layer 3 — Report signups, not sessions

Add signup as a key event, then read it against first-user scope as well as session scope. GA4 key events use data-driven attribution by default with a 90-day conversion window for non-acquisition events, so an AI first touch followed by a branded-search return can still earn partial credit — but only in the attribution reports, not the session-scoped Traffic acquisition table.

Layer 4 — Close the loop to CRM

BigQuery export lets you join session data to CRM outcomes on user_id; it's free but capped at 1 million events per day on standard properties, and the streaming export excludes new-user and new-session traffic source data — check which export you're on before building the join. Measurement Protocol pushes offline stages (SQL, closed-won) back into GA4; it requires client_id, plus session_id sent within 24 hours of session start to bind the event to that session. Google warns in the same docs that "only partial reporting may be available."

Layer 5 — A self-reported attribution field

Add "How did you hear about us?" to signup and demo forms with an explicit ChatGPT / AI assistant option. It is the only instrument that catches the ~56% of AI-influenced visits arriving via branded search. Self-reporting is noisy, so treat it as a cross-check on magnitude rather than a replacement for GA4.

Layer 6 — Cross-check against first-party citation data

Bing Webmaster Tools' AI Performance report, released 11 February 2026, is a free first-party view of which pages an answer engine actually used: grounding queries (the phrases Copilot generates internally when it retrieves content) and citations (how many times Copilot used a given page). Otterly published three months of its own as an example — 647 grounding queries, 30,398 grounding events, 173 pages. Google also added generative AI performance reports to Search Console in June 2026; check what your property exposes.

Five mistakes that cost the most

  1. Filtering on medium instead of source, which kills most ChatGPT sessions.
  2. Leaving the AI rule below Referral in the ordering.
  3. Presenting the native channel's start date as a growth curve.
  4. Counting AI Overviews traffic as an AI channel — by Google's definition it's Organic Search, so double-counting is easy.
  5. Assuming every (not set) is an AI attribution problem. Consent mode, tag firing order, server-side transformations and Measurement Protocol misuse all produce Unassigned and (not set) too.

Joining the visibility score to the GA4 report

GA4 can tell you an AI-assistant session landed on /pricing. It cannot tell you which prompt produced it. That side of the join only exists inside a visibility tool.

The join looks like this:

tracked prompt → engine that answered it → URL cited → GA4 landing page → key event

Run it as a table, not a dashboard. For each buyer-intent prompt, record which engines cite you and which URLs they cite, then pull sessions and key events for those landing pages from your custom AI channel. Four-engine coverage earns its place here: Claude and Perplexity aren't in GA4's native channel and Google's AI Mode never appears as an AI channel at all, so a visibility score is the only read you get on those surfaces. Geoptimizer's per-prompt, per-engine citation records are built for this table — the in-app report history keeps every scan's per-engine breakdown next to your GA4 pull.

Three rules keep the join from lying to you:

Lag the comparison by two to six weeks. AI answers favor recently-updated pages: Seer's citation study found 75% of pages LLMs cite were updated within the last year, and the pages cited across all four study months were established-and-maintained rather than freshly published. Correlating this week's score against this week's sessions will under-detect a real effect.

Use rolling windows, not single scans. AI answers are nondeterministic, so a headline score belongs in a rolling window with a confidence band and a single on-demand scan is a snapshot. That discipline also stops you misreading a frontier-model release as a campaign result — our walkthrough on telling a real visibility shift from model-release noise covers the diagnostic. Same rule on the GA4 side: 28- or 90-day rolling windows.

Compare share of signups, not conversion rate, at low volume. With AI at roughly 1–3% of sessions for most B2B sites, a monthly AI-channel conversion rate is mostly noise. One widely shared comparison — AI Search 20.7% versus Direct 11.1% — rests on 87 AI sessions. Interesting; not a benchmark. And if two GEO tools hand you different inputs for this join, our guide to why AI visibility tools report different scores is the audit for deciding which to trust.

An illustrative model

Built from published benchmarks, not one company's data — label it that way when you present it. For a B2B SaaS site at 100,000 monthly sessions:

  • AI-assistant sessions at ~1.5% of total = 1,500/month visible. (Previsible's panel of 166 GA4 properties and 6.77M LLM sessions puts most verticals near that band, with insurance highest at 1.51% of sessions.)
  • Apply Clickport's 35.7% referrer-less share and the true figure is closer to 2,300/month; apply Loamly's 70.6% and you'd get ~5,100, but that upper bound leans on the lowest-confidence detection layer. Recovering the UTM-only sessions with Layer 1 and 2 moves you toward the lower number.
  • Apply HockeyStack's median session→hand-raiser rate of 5.24% → ~120 hand-raisers/month, of which 86% are high-intent.
  • Apply the median hand-raiser→pipeline rate of 2.66% → ~3 opportunities/month.
  • Then add the downstream layer: if ~56% of AI-influenced visits arrive via branded search, the AI channel is capturing well under half the demand it created.

What to actually put in front of a CFO

Three numbers, plus one experiment.

1. Share of new signups whose first touch was an AI assistant. Share of signups survives small samples in a way conversion rate does not. This is the Ahrefs argument — 0.5% of visits, 12.1% of signups — and the Webflow one: LLM-sourced signups went from 4% of all signups in Q1 2025 to 8% in Q2 2025, at 24% signup conversion versus 4% for non-brand SEO. By late 2025, VP of Growth Josh Grant put it at 10% of signups from AI discovery, growing 4x year over year, with two in three LLM referrals converting within 7 days. His framing is the CFO-ready one: aggregate traffic fell while quality traffic rose, so volume hid the win.

2. AI-influenced pipeline, deduplicated. First-touch AI sessions plus self-reported "heard about you via ChatGPT," overlaps removed. Report pipeline created, and closed-won as a range once you have enough of it — see the caveats below for why.

3. Branded search and direct trend as the non-click proxy. Branded impressions and direct sessions, plotted against your visibility score with the lag applied. This is the only line on the page representing the 56%.

The experiment: a paired page-cohort test. Attribution can't answer incrementality, so run a test instead. Optimize half of a set of comparable pages for AI citation, hold the other half as a control, state the window up front, and compare citation gains and landing-page key events across cohorts. Graphite ran control-group content experiments with t-tests (p = 0.001 and 0.002) in the Webflow engagement — that's the standard to aim at.

The caveats that keep you credible

A GEO business case dies faster from one over-claim than from a modest number.

For practical methods to break visibility down by market and compare language- and country-level results, see AI visibility by country and language: Measure per market.

Pipeline follow-through is genuinely weak for most B2B companies right now. HockeyStack Labs' analysis of 118 accounts found a median session→hand-raiser rate of 5.24%, hand-raiser→pipeline of 2.66%, and median pipeline→closed-won of 0%. Only the 501–1,000 employee segment performed well (~17% session→hand-raiser, 5.6% into pipeline, 2.8% to closed-won), and the report doesn't disclose how LLM sessions were identified. Say this out loud: today, citations reliably turn into hand-raisers, and closed-won evidence sits in one size band.

The conversion-multiple literature disagrees with itself, and scope is why. 23x (Ahrefs' own site, 30 days), 4.4x visitor value (Semrush, 500+ marketing/SEO topics, explicitly extrapolation-heavy), ~6x (Webflow signups), +42% conversion (Adobe, North American retail, March 2026 — vendor-stated, methodology not public, measured against all non-AI traffic combined), and from the other direction a study of 973 ecommerce sites comparing 50,000+ ChatGPT-driven transactions against 164 million from other channels that found affiliate converting 86% higher than ChatGPT. Anyone quoting one multiple as a law is misreading the other four. The defensible claim is directional — AI-referred visitors tend to arrive later in the research process and convert better than the same site's non-brand organic — and magnitude has to be measured locally.

For a deeper look at the 4.4x claim, the methods behind these published multiples, and how different definitions change the headline number, see our AI referral traffic conversion rate analysis.

Incrementality is the strongest objection, not measurement. Rand Fishkin asked it plainly of the Similarweb data: "Would these users have found these brands anyway?" The brand-level lifts there are far more modest than the headline 2.5x — American Express +7.2%, Capital One +14.2% more likely to receive a visit when recommended, with traditional search traffic down around 15% — and the study doesn't disclose panel size or dates.

Some questions have no public answer yet. Previsible states it directly: "Conversion rate by LLM platform is the single most valuable unanswered question in AI discovery." Nobody has published a clean study correlating a third-party visibility score with same-site referral traffic over time. If a vendor offers you one, ask for the method.

And one tactic to avoid. Pre-tagging your own canonical URLs with UTMs so assistants pass them through pollutes those URLs, risks self-referrals and duplicate landing-page rows, and can leak into organic results. Tag only links you place off your own site.

Why the math is moving in your favor

Build this now because the exposure-to-click rate just changed. On 7 May 2026, ChatGPT began embedding clickable brand links inside answers. Qwairy's analysis of more than 140,000 ChatGPT answers collected 1 April–21 May 2026 found the share carrying one jumped from roughly 0.4% to 6.2% in a single day, with homepage destinations rising from 59% to 79% of brand links — every post-shift link tagged utm_source=chatgpt.com. The authors are careful: this measures links written, not clicks, and the rate eased from a 7–10% first-week peak toward roughly 4–5% by 21 May. Similarweb measured the traffic side at +157.7% ChatGPT referrals week over week. And SE Ranking's dataset of 101,574 websites found 60% of AI-referred traffic lands on homepages versus 17% for organic search — so check your homepage, not just your blog, when you look for the effect.

The gap between "mentioned" and "clicked" is narrowing. Your last GA4 channel audit is probably older than the change that narrowed it.

FAQ

Why is my ChatGPT traffic showing up as Direct or Unassigned in GA4? Because ChatGPT appends utm_source=chatgpt.com without a utm_medium, and GA4 classifies on UTM values whenever any UTM parameter is present, ignoring the referrer. With a source but no medium, the session matches no channel rule and falls into Unassigned. Separately, rel="noreferrer" links and in-app browsers send no referrer at all, which lands them in Direct. Building your AI channel rule on session source rather than medium recovers the first group.

Does GA4's AI Assistants channel track Claude and Perplexity? No. Google's documented list names ChatGPT, Gemini, DeepSeek, Copilot and Grok. Claude and Perplexity need their own rules in a custom channel group, keyed on source. Note that Google has already revised this list once, so re-check it quarterly.

Can I see AI Overviews traffic separately in GA4? No. Google defines Organic Search as including AI Overviews and AI Mode, so those clicks are deliberately blended into Organic Search and no configuration separates them. Search Console's generative AI performance reports and your branded-search trend are the practical proxies.

For methods to measure visibility when clicks are absent, consult Zero-click AI answers: measuring visibility without traffic.

How long before a visibility score change shows up in GA4? Plan for two to six weeks and lag your comparison accordingly. Citation sets churn, and the pages with staying power tend to be established ones that get maintained rather than newly published — so a content change takes weeks to settle into a stable traffic effect.

What's the minimum sample before I report an AI conversion rate? Higher than most published examples: one frequently shared comparison rests on 87 sessions. Below a few hundred AI sessions in the window, report share of signups and a rolling 28- or 90-day trend instead, and say plainly that the AI number is a directional floor.

Bringing it together

The two numbers were never going to reconcile perfectly, and a plan that depends on them doing so fails its first review. What works is narrower: fix the classification so GA4 stops mislabeling the AI sessions it can see, report share of signups while volumes are small, add a self-reported field to catch the branded-search majority, and run a paired page-cohort test so the incrementality answer is ready before anyone asks. Pair that with a visibility score covering the engines GA4 can't — and be candid about the caveats, because the caveats are what make the number believable.

The left-hand side of that join is the part GA4 will never give you: which prompt, which engine, which cited URL. Run a free four-engine visibility scan — ChatGPT, Gemini, Claude and Grok are included on every plan, including the free one — then map your cited URLs to your GA4 landing pages and see how much of your AI-influenced demand is currently filed under Direct.

Keep reading

See it on your own domain.

Free visibility check across ChatGPT, Gemini, Claude, and Grok — about 30 seconds.

Run the free check