· 18 min read · Geoptimizer Team
AI Crawler Log Analysis: Crawl Volume Isn't Citations
- technical-geo
- ai-crawlers
- log-analysis
- gptbot
- measurement
Server logs answer one question well: can these bots reach my pages? They cannot answer do these engines use my pages in answers? — and the gap between those two questions is where most AI crawler reports go wrong. GPTBot is a training crawler that, by OpenAI's own documentation, has nothing to do with ChatGPT search answers, and even when the right retrieval bot does fetch a page, studies put the odds of that fetch becoming a visible citation somewhere between roughly half (Ahrefs, 1.4 million ChatGPT prompts) and about one in seven (AirOps, 548,534 retrieved pages). A defensible report keeps three layers separate: access from logs, citation from prompt-level answer tracking, referral from analytics.
The report that says everything and proves nothing
Here is the scenario that keeps producing confident, wrong slides. Someone asks you to prove the AI bots can reach the site. You pull thirty days of logs, filter on GPTBot, find 4,000 page fetches with a clean 200 rate and no blocked responses, and write "AI crawlers have full access, 4,000 pages crawled last month" — accurate, checkable, first-party, unsampled. Then the brand lead points out the company has never once been cited in a ChatGPT answer, and asks what happened.
Nothing happened. The two facts aren't in conflict, because they aren't about the same thing. OpenAI's crawler documentation describes GPTBot as the bot that crawls "content that may be used in training our generative AI foundation models," and the only consequence it attaches to blocking GPTBot is that your content "should not be used in training." The sentence that matters for answers belongs to a different bot: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." GPTBot isn't in it at all.
So a month of heavy GPTBot traffic tells you your robots.txt lets a training crawler in and your origin served it bytes. That's worth knowing — it's the floor of AI visibility, and a broken floor is fatal. But it is a floor, not a ceiling, and reporting it as a visibility trend is how a technical SEO ends up defending a number that mostly describes a training bill.
What each user agent is actually evidence of
Start with the taxonomy, because everything downstream depends on it. Every major operator that publishes bot documentation now runs a three-way split — training crawler, search-index crawler, user-initiated fetcher — and the three carry completely different implications.
| Operator | Training | Search / index | User-initiated |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | none (Perplexity states no training use) | PerplexityBot | Perplexity-User |
Anthropic's crawler documentation draws the same line, and is explicit that blocking Claude-SearchBot "reduces site visibility in user search results" — a consequence it doesn't attach to ClaudeBot. Perplexity's docs describe PerplexityBot as "designed to surface and link websites in search results on Perplexity," and say neither bot feeds model training.
Two details in that table will bite you. First, user-initiated fetchers aren't bound by robots.txt the way crawlers are: OpenAI states that "because these actions are initiated by a user, robots.txt rules may not apply," and Perplexity says Perplexity-User "generally ignores robots.txt rules." A ChatGPT-User hit proves a user asked about something on your domain — not that your robots.txt is permissive, which is exactly the meaning people give it.
Second, and this is the misread that shows up most often: there is no Gemini crawler to find. Google's AI features documentation states there are "no additional requirements to appear in AI Overviews or AI Mode," because "AI is built into Search… which is why robots.txt directives for Googlebot is the control." Google-Extended is a robots.txt token rather than a user agent, governing whether crawled content may be used for Gemini training and grounding, and Google notes it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." So a flat line for Google in your AI bot report isn't evidence of a Gemini problem — there is nothing to see, and a chart implying otherwise sends a team off to fix a crawler that doesn't exist.
The reframing that fixes most of this: logs answer can they read me, per bot and per purpose. Everything in a log report should fit inside that sentence.
Why volume is the wrong axis
Once the bots are separated by purpose, the next instinct is to plot total hits over time — and that's where the second layer of self-deception lives, because most AI crawl volume on the open web is training volume.
Cloudflare's network-wide breakdown for July 1–28, 2025 put training at nearly 80% of AI bot crawling, with "user action" and "undeclared" together under 5%. Its one-year-on report puts training at 52% of crawler requests as of June 2026, up from 22% in Spring 2025, with mixed-use crawlers over 36% and pure search "a small and declining share of overall crawler activity." Those shares look contradictory and aren't — Cloudflare's taxonomy changed, adding a mixed bucket that absorbs multi-purpose crawlers — so quote either with its date. The direction holds in both, and that alone should stop you plotting one undifferentiated "AI bots" line.
The weak link between crawl and anything downstream shows up more starkly in Cloudflare's crawl-to-refer ratio, which divides a platform's HTML requests by the requests carrying its referrer. In the first week of August 2025, Anthropic sat at roughly 50,000:1 network-wide, OpenAI at 887:1, Perplexity at 118:1. Tens of thousands of crawled pages per referral is normal, not pathological. And the spread inside one operator kills cross-industry benchmarking: OpenAI's ratio that same week was 152:1 in News & Publications but 401.7:1 in Computer & Electronics. Benchmark against a whole-web average and you can report a healthy-looking ratio that's several times worse than your actual peer set.
Then there's the trend problem. Botify analysed roughly 7 billion OpenAI bot log events from November 2024 to March 2026 and found OpenAI's crawling tripled after GPT-5's August 2025 launch — OAI-SearchBot up 3.5x, GPTBot up 2.9x. In the same dataset, ChatGPT-User activity fell 28% comparing December 1 to March 14, 2026 against the prior equivalent period. The bot closest to a live user question went down while the total went up; a volume-only chart would have called that quarter unambiguous growth. (Botify's corpus skews to large enterprise sites, as Search Engine Journal noted.) Which raises the obvious question: when the right bot fetches a page, what are the odds it turns into anything?
The two gaps that break the chain
There are exactly two places the chain from log line to citation snaps, and they break in opposite directions.
Gap one: retrieval happens, citation doesn't. Ahrefs analysed 1.4 million ChatGPT prompts on GPT-5.2 desktop (published April 2026) and found the assistant retrieves dozens of URLs per query and cites about half — 49.98% cited, 50.02% not. AirOps, looking at 548,534 pages retrieved across 15,000 prompts, found only 15% appeared in final answers, yielding 82,108 citations. The numbers disagree because the denominators do: Ahrefs counts URLs the client retrieved, AirOps counts pages discovered during research. Present them as a range with their definitions rather than averaging into fake precision. Either way, being fetched by a retrieval bot is a lottery ticket — somewhere between a coin flip and one in seven.
That changes what a fetch is worth in a report. "OAI-SearchBot fetched 1,200 of our URLs last month" honestly translates to "we bought between 180 and 600 tickets," and the useful follow-up is which pages tend to win. Ahrefs found citation rate varied enormously by retrieval channel — 88.46% for the general search channel versus 1.93% for Reddit — while AirOps found 55.8% of cited pages ranked in Google's top 20, with position-1 pages cited 3.5x more often than pages outside it. None of those levers appear anywhere in a log file.
Gap two: citation happens, the fetch doesn't. This is the one that quietly invalidates "we saw no hits, so they can't be using us." Seer Interactive matched SearchGPT citations against SERPs for 100 queries and 500+ citations and found 87%+ matched Bing's top organic results versus 56% for Google — a sample they call directional, but directionally clear. And Aleyda Solís ran a test starting July 24, 2025 in which ChatGPT, asked about a brand-new page, produced an answer matching Google's SERP snippet word for word and called its own source "a cached snippet via web search." The retrieval landed on a search index, not your origin — no log line exists for it on any server you control. It happens at publisher scale too: Digiday reported that The New York Times blocks the ChatGPT and Perplexity crawlers in robots.txt and still received 240,600 visits from ChatGPT in January 2025.
Access still matters, and it would be dishonest to imply otherwise. cloro.dev studied 1,058 prominent domains and found a median citation propensity of 0.003 citations per Google-organic appearance for domains blocking GPTBot versus 0.417 for domains allowing it — a gap the authors say "can't prove causation." The same study found publishers drawing a deliberate line: 13.9% block GPTBot at the root, only 3.4% block OAI-SearchBot. "Don't train on me, do cite me" is a coherent policy, not a mistake. Access is the floor; it isn't the ceiling.
Six checks that make a log report defensible
None of this makes logs useless. It makes a raw hit count useless. Here's the sequence that turns one into something you can put your name on.
1. Verify before you count. Match source IPs against the operator's published ranges — OpenAI publishes a separate JSON per bot (gptbot.json, searchbot.json, chatgpt-user.json), Anthropic one at claude.com/crawling/bots.json, Perplexity per-bot files — or accept your CDN's verified-bot classification. Cloudflare has folded HTTP Message Signatures into its Verified Bots Program, validating ed25519 signatures at the edge. In Screaming Frog's Log File Analyser, "Verify Bots When Importing Logs" "performs a lookup against publicly confirmed IP lists to confirm they are genuine." Never ship a user-agent-only count.
Why that isn't paranoia: on unindexed honeypot domains whose robots.txt disallowed all automation, Cloudflare observed an undeclared crawler at 3–6 million daily requests presenting a generic Mozilla/5.0 (Macintosh…) string from IPs outside the published range. The same post records the contrast — "ChatGPT-User fetched the robots file and stopped crawling when it was disallowed." User agents are self-reported strings, and compliance varies by operator.
2. Split by purpose, never by vendor. "OpenAI hit us 40,000 times" is not a finding. GPTBot belongs on a training line; OAI-SearchBot and ChatGPT-User belong on a retrieval line that can plausibly connect to a citation. One row per vendor answers a question nobody asked.
3. Count unique URLs and coverage, not requests. Request totals are dominated by retries, redirects and dead URLs. Vercel and MERJ found 34.82% of ChatGPT's fetches and 34.16% of Claude's returned 404s across their network, against 8.22% for Googlebot — so if a third of a bot's requests hit pages that don't exist, your crawl-activity chart is substantially a measure of that bot's URL hygiene. Bursts compound it: one small CDN study logged GPTBot making 187 requests in a week with 152 inside a single three-minute window — an n=1 site, but a good picture of how "4,000 pages crawled" can be one queue draining on a Saturday. The reportable metric: what share of your sitemap URLs were fetched with a 200 by a retrieval bot in the last 30 days, and how recently.
4. Report the status-code distribution per bot. This is the part of "prove the AI bots can reach us" that logs genuinely answer. robots.txt records intent; only a log shows a 403 or a managed challenge being served to a verified retrieval bot.
5. Exclude the noise rows — and remember a 200 isn't a read. Strip robots.txt, sitemap.xml, favicons and static assets before totalling; they behave very differently per bot (in that same study, OAI-SearchBot fetched robots.txt around 3.8 times a day while GPTBot never fetched it across 48 days). Watch for asset padding too — images were 35.17% of Claude's fetches on nextjs.org in Vercel's data. Then the harder caveat: none of the major AI crawlers in Vercel's data executed JavaScript. On a client-rendered page, "GPTBot fetched it successfully" and "GPTBot read nothing" produce the identical log line.
6. Know what your log source cannot see. Origin logs miss everything served from CDN cache. Cloudflare Logpush — raw per-request lines — is Enterprise-only; on Free, Pro and Business you have dashboard aggregates, which is the wall many people hit the moment they're told to "just check the logs." AI Crawl Control fills part of that gap on all plans, showing which AI services access your content and which crawlers follow your directives.
No per-request logs at all? Do the cheap version first: run the domain through Geoptimizer's free AI Crawler Checker, which reports whether GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and six others can reach your site, then work through the 20-minute crawler access and llms.txt setup. Configuration errors are cheaper to find than log anomalies, and they explain most access failures.
What logs will never tell you — and what now does
No log anywhere records an AI Overview, a Gemini answer, or a ChatGPT answer built from an index snippet. The good news is that the citation side stopped being a black box in 2026, as three first-party sources shipped — each covering a slice, none substituting for the others.
Bing Webmaster Tools launched an AI Performance report in public preview on February 10, 2026: total citations, average cited pages, page-level counts, and "grounding queries" — the phrases AI used when retrieving referenced content — across Microsoft Copilot, Bing's AI summaries and select partners. Microsoft states the data "does not indicate ranking, authority, or the role of any page within an individual answer." Given Seer's Bing-overlap finding, it's the closest free proxy for ChatGPT-adjacent retrieval — but it covers Microsoft surfaces only.
Google Search Console added generative AI performance reports on June 3, 2026, to a subset of properties. They report impressions from AI Overviews and AI Mode by page, country, device and date — and currently no queries, no clicks, no CTR, no position, no citation placement. Useful, and much less than people assume when they see the tab.
GA4 now has an AI Assistants default channel for users arriving from ChatGPT, Gemini, Copilot or Grok. That's referral, not citation, and page-level attribution is degrading: Similarweb reports that after ChatGPT's May 2026 homepage-link change, homepage landings jumped from 26–29% to 62–63% of its referrals.
Which leaves the engines that publish nothing. There's no Gemini crawler and no Gemini citation report; Claude offers neither; and Grok's answers draw partly on X search, a citation source with no crawl signature on your domain at all. For those, the only instrument is prompt-level answer tracking: run the questions your buyers actually ask, on each engine, with web search on, and count how often you're mentioned and cited. That's what Geoptimizer does — buyer-intent prompts across ChatGPT, Gemini, Claude and Grok, scored on mention rate, citation rate, prominence and sentiment with a published, versioned formula. Citation rate carries 25% of each engine score, the natural counterpart to the coverage number from your logs.
| Question | Instrument | What it can't tell you |
|---|---|---|
| Can the bots reach my pages? | Server / CDN logs, verified by IP or signature | Whether anything was cited |
| Was I cited on Microsoft surfaces? | Bing WMT AI Performance | ChatGPT, Gemini, Claude, Grok |
| Was I shown in Google's AI surfaces? | GSC generative AI reports (impressions only) | Queries, clicks, placement |
| Did anyone click through? | GA4 AI Assistants channel | Answers with no link click |
| Am I mentioned and cited across engines? | Prompt-level answer tracking | Whether a bot ever fetched the page |
Annotate your chart: the September 15 discontinuity
One last thing before you ship a year-over-year AI bot chart. Cloudflare has replaced the single "block AI bots" toggle with three purpose categories — Search, Agent and Training, each settable to allow, block everywhere, or block only on pages displaying ads, available to all customers including Free since July 1, 2026. From September 15, 2026, new domains get new defaults: Training and Agent blocked on ad-supported pages, Search still allowed, multi-purpose crawlers blocked for their training component. Existing domains are unaffected unless they opt in.
The reporting consequence is easy to forget: any AI bot chart crossing September 15, 2026 contains a policy discontinuity, so some share of a Q4 change will be a default change rather than an engine change. Annotate the date and note whether your zone opted in. If you're still deciding which categories to allow, our bot-by-bot framework for Cloudflare's September 15 defaults covers that decision in full.
The one-page report that holds up
Put the two halves side by side and label them, because the fastest way to lose a stakeholder's trust is to let them reconcile columns that were never going to agree.
Access metrics (verified logs, per bot, last 30 days): verified requests; unique URLs fetched with a 200; sitemap coverage percentage; non-200 rate by status class; median days since the last retrieval-bot fetch.
Answer metrics (prompt-level tracking, per engine, rolling window): mention rate; citation rate; prominence; share of tracked prompts where a competitor appears and you don't — each with a confidence band.
Then one line at the top: these measure different things and will not move together.
That's the finding, not a hedge. A July 2026 variance-decomposition study of 12,933 LLM responses across 20 brands, 8 languages and 3 models found brand identity explained just 1.5% of response-level variance, with 34.8% coming from pure resampling, and brand-ranking reliability rising from 0.01 for a single response to 0.36 across the full design. That is the strongest form of the objection technical SEOs raise here — logs are the only ground truth I have; the citation trackers are the flaky ones — and it's half fair. Logs are precise about the wrong question; citation tracking is noisy about the right one. So run both, label single scans as snapshots, and report the answer side as a window with a band. For separating a genuine shift from that noise, our diagnostic for model releases versus measurement noise walks through the method.
FAQ
Does blocking GPTBot hurt my ChatGPT visibility? Per OpenAI's documentation, blocking GPTBot means your content "should not be used in training generative AI foundation models." The bot tied to answers is OAI-SearchBot, and opting out of that one means sites "will not be shown in ChatGPT search answers, though can still appear as navigational links." For most brands the low-risk configuration is to allow the retrieval bots and treat training as a separate business decision.
Why do I see no Gemini crawler in my logs? Because there isn't one. Google states there are "no additional requirements to appear in AI Overviews or AI Mode" and that robots.txt directives for Googlebot are the control. Google-Extended is a robots.txt token governing Gemini training and grounding, not a user agent in your access logs, and Google says it doesn't affect Search inclusion or ranking.
How many crawled pages should turn into citations? There's no reliable conversion rate, and that's the point. Of pages the retrieval layer actually pulls, Ahrefs measured 49.98% cited and AirOps 15% — different denominators, both far below 100%. For training crawls like GPTBot's there's no measurable short-term link at all; training feeds models that ship later with published knowledge cutoffs, so any effect is lagged and unattributable in a quarterly report.
Can an AI engine cite my page without ever fetching it? Yes. Aleyda Solís documented ChatGPT answering about a new page using text matching Google's SERP snippet verbatim, describing its source as "a cached snippet via web search." Seer Interactive found 87%+ of SearchGPT citations matched Bing's top organic results. Engines can answer from an index snippet, a partner index, or model memory — none of which touches your origin.
I'm on Cloudflare's free plan. Can I still do this? Partly. Logpush raw per-request logs are Enterprise-only, but AI Crawl Control works on all plans and shows which AI services access your content and which crawlers follow your directives, and origin logs still capture anything that misses cache.
Before your next AI crawler report
Logs answer can they read me, per bot and per purpose, and nothing beyond that. Citation tracking answers do they use me, noisily, which is why it needs a rolling window and a confidence band. Report the first as if it were the second and you'll eventually be asked to explain 4,000 crawled pages and zero citations.
Fix the access floor first, then measure the ceiling. Confirm the retrieval bots can reach you with the free crawler check above, then start tracking what the engines actually say — Geoptimizer's free plan covers all four engines with no per-engine add-ons, so you can have a real citation number sitting next to your log coverage number before the next report is due.