· 18 min read · Geoptimizer Team
How Often Does ChatGPT Search the Web? About a Third
- generative-engine-optimization
- ai-visibility
- measurement
- chatgpt
- retrieval
ChatGPT enabled web search on 34.5% of queries in February 2026, down from 46% in late 2024, according to Semrush's analysis of more than a billion lines of US clickstream data from its 200-million-user panel. The other two-thirds of the time, the model answers from what it already holds in its weights — no fetch, no citation, no page of yours involved. That split decides whether your content roadmap can move your AI visibility at all: retrieved answers respond to pages, crawler access and third-party lists, while remembered answers respond only to how well your brand was represented across the web before the model's knowledge cutoff. If you publish constantly and your mentions haven't budged, you are probably buying the first fix for the second problem.
The stakes stopped being theoretical for B2B software this year. A G2 survey of 1,076 buyers and decision-makers conducted in March 2026 found 51% now start research with an AI chatbot more often than with Google, up from 29% in April 2025, and 69% said AI guidance led them to choose a different vendor than the one they had planned on. As G2's Tim Sanders put it, "buyers have moved from reference to inference." Which turns a mechanical question — does the engine even look at the live web when your buyer asks — into a pipeline question.
Two engines behind one answer box
There is no ambiguity about how this works, because OpenAI documents it. The web search tool is a tool, and the API docs state plainly: "Like any other tool, the model can choose to search the web or not based on the content of the input prompt." You can force it with tool_choice: "required" or leave it on auto. Anthropic describes the same behaviour in Claude's web search documentation: Claude "searches when the request depends on information that is current, changing, or outside its training data" and "answers directly without searching when the request draws on stable knowledge." Both vendors are saying the same thing. Retrieval is conditional. Memory is the default.
That was the design intent from launch: OpenAI's product lead for search, Adam Fry, told MIT Technology Review in October 2024 that "searches about generalized topics will still draw on this information from the model itself," while recent things — sports, stocks, news of the day — trigger an automatic search.
The cleanest proof that these are two separate systems is that OpenAI runs separate crawlers for them. GPTBot collects training data; disallowing it signals your content shouldn't be used to train foundation models. OAI-SearchBot feeds search answers, and the docs are explicit that sites opted out of it "will not be shown in ChatGPT search answers." ChatGPT-User handles live, user-initiated fetches and isn't used for automatic crawling. Three user agents, three jobs, three different consequences. Your robots.txt is not a privacy setting; it is a routing decision about which half of your visibility you keep — and changes take roughly 24 hours to register, so a block set months ago is still shaping today's answers.
Understanding the split is the easy part. The harder question is which half your buyers' prompts land in, and here the headline number quietly misleads.
"About a third" is the wrong number for your prompt set
The 34.5% figure is a population average across everything anyone types into ChatGPT, and much of that population was never a candidate for search. Profound's intent study, drawn from tens of millions of real ChatGPT prompts, found the mix breaks down as generative 37.5%, informational 32.7%, no intent 12.1%, commercial 9.5%, transactional 6.1% and navigational 2.1%. Nearly four in ten prompts are "write me this email" or "rewrite this paragraph." They will never search, and they sit in the denominator of every "only a third" statistic you have read.
Strip them out and the picture changes. In an observational study of AI citation behaviour published in April 2026, 391 brand and product queries run through the ChatGPT web UI triggered a search 42% of the time; in the same researcher's API dataset, the rate ran to roughly 73% for discovery queries, 70% for review-seeking and 65% for comparison, against about 10% for purely informational ones. Nectiv's AI Tracker points the same way, finding 31% of 8,500+ prompts triggered at least one search with 2.17 searches per searching prompt, and a vertical spread from 59% on local intent down to 18% on credit cards. As Nectiv's Chris Long put it, "when ChatGPT uses search, SEOs have much more control over the information that's presented."
So the number that matters isn't a third — it's yours. Two teams tracking fifty prompts each can genuinely sit at 15% and 70% depending on whether their prompts read like "best CRM for field service teams" or "what is a customer data platform." That is the missing column in most tracking sheets: not just did the engine mention us, but did it search at all.
Three complications are worth building into your expectations before you start logging it.
Search fires early and fades. Profound's analysis of roughly 730,000 ChatGPT conversations containing at least one web citation found citations appear in 12.6% of first turns and drop to around 3% by turn 20 — turn one is about 2.5x more likely to cite than turn ten. Your tracked prompts are nearly all turn-one prompts, so your dashboard already samples the most retrieval-friendly slice of real usage. A buyer five messages into a conversation is seeing more of the memory half than you are.
Model and plan tier change the answer. RESONEO's tracking of 400 prompts and around 27,000 responses over 14 weeks found GPT-5.3 Instant runs two to three visible search turns while GPT-5.4 Thinking runs five to ten or more, and reported free-tier prompts falling back to smaller models with no web search tools at all. Treat that as one lab's observation rather than documented behaviour — but weigh it against the fact that, as Search Engine Land notes, over 90% of ChatGPT's roughly 900 million weekly users are on free plans.
Nobody has ground truth. OpenAI publishes no trigger-rate data. Clickstream panels infer search from observable signals, citation-based methods only see conversations where a citation rendered, and UI tests use small hand-built prompt sets. Semrush's own series swings between 15% and 66.3% across the study period — note it switches units there, from queries to sessions — while Profound's conversation-level read lands near 18%, and Otterly, another GEO tool, publishes a 20–35% range alongside a candid note that "OpenAI does not disclose any specific data regarding how frequently the web search feature is triggered." The honest read: roughly a third, wide error bars, and falling every month from November 2025 to February 2026.
Why your new pages aren't moving your mentions
If retrieval fires on most of your commercial prompts, publishing ought to work. So why doesn't it show up?
Rule out the explanation most teams reach for first. Profound tracked around 900 newly published marketing pages over a 60-day window in March–May 2026 and found the median time from publication to first citation on ChatGPT or Claude was 6.81 days, with P75 at 18.68 days and P90 at 37.10 days. Half get cited inside a week, ninety percent inside about a month. "The engines haven't found my pages yet" is almost never true after a quarter. They found them. They just didn't need them.
What they needed was somebody else's ranked list. An analysis of roughly 25,000 unique URLs and around 400 million citations across six AI surfaces in March–April 2026 found 63% of all citations pointed to listicles, with ranked lists making up 71–86% of those. In B2B SaaS the skew is sharper. Overthink Group ran 1,263 solution-aware prompts across ChatGPT, Gemini, Perplexity and AI Overviews in one business week in June 2026, covering 250 niche categories, and found 70.8% of all citations were "best" or "top" listicles, 51.6% of them carrying "2026" in the title. The G2 network took 8% of citations; Reddit took 1.4%. And the engine you optimise for matters: Google's AI surfaces cited actual software vendors in roughly 75% of their top 100 domains, ChatGPT and Perplexity in only about 25%.
The starkest version comes from Victorious's Q2 2026 study, which found that within category-research prompts, 99.99% of the 49,391 citations analysed pointed to third-party websites rather than the brand's own domain — only four of 150 brands earned a citation to their own site. For a content lead that reframes the roadmap: on category prompts, your realistic best case usually isn't your page being cited, it's your product being named inside a page you don't own.
None of which makes page-level work pointless. Ahrefs' analysis of a 1.4-million-prompt ChatGPT dataset found about half of retrieved URLs end up cited, and that cited URLs matched the model's fan-out sub-queries more closely (0.656 cosine similarity) than the original prompt (0.602) — a concrete instruction to write for the sub-questions an engine decomposes a prompt into, not the prompt you pasted into your tracker. Natural-language slugs were cited 89.78% of the time versus 81.11% for opaque ones. The academic baseline agrees: the KDD 2024 "Generative Engine Optimization" paper measured up to 40% visibility lift from source-page tactics like adding statistics, quotations and citations. And one familiar lever still dominates — in that April 2026 study, URLs ranked first on Google were cited by at least one AI platform 54% of the time, versus roughly 2% at position 100.
So retrieval-side work pays, when retrieval happens. The problem is the prompts where it doesn't.
The memory half: recognised, but never recommended
The most uncomfortable number in this topic comes from that same Victorious report, which tested 175 brands across legal, healthcare, SaaS, financial services and ecommerce on eight AI platforms. Those platforms accurately described 96% of brands when asked about them directly, yet 89% of the brands never appeared in answers to category-research questions. The models know who you are. They just don't bring you up. As the report puts it, "recognition and mention are distinct, measurable behaviors" — so if your testing only asks "does ChatGPT describe us correctly," you are passing an exam that has nothing to do with pipeline.
The same study sketches a rough floor: brands with fewer than 2,000 indexed web pages mentioning them were named in AI answers just 3% of the time. It also maps the journey gap — at problem-awareness stage brands were named in only 0.10% of answers, while category-research prompts named them more than twelve times as often. If your prompt set leans early-stage, your score is measuring a slot that barely exists for anybody.
What correlates with living in that memory isn't what most SEO teams have budgeted for. Ahrefs' study of 75,000 brands found Spearman correlations with ChatGPT visibility of 0.737 for YouTube mentions and 0.664 for branded web mentions, against 0.266 for Domain Rating and roughly 0.19 for number of backlinks — mentions out-correlating links by more than three to one. Ahrefs also says, in its own words, "correlation isn't causation," and the obvious confound is brand size, since large brands have more of everything. Read it as a directional signal about where unowned surface area lives, not a dose-response curve.
Timing is what makes memory work feel thankless. GPT-5.6's Sol, Terra and Luna models all carry a stated knowledge cutoff of 16 February 2026 — roughly six months behind the calendar. Everything published after mid-February exists for those models only through retrieval. A PR push landing this month can't reach the weights until a future generation trains on a web that contains it. That's a quarters-long feedback loop, and no dashboard will show progress on it in week two.
One stronger claim deserves careful handling. RESONEO's reverse-engineering of ChatGPT's fan-out behaviour concluded that "a brand absent from parametric memory won't even be considered as a search candidate" — memory gating retrieval, rather than sitting beside it. It's one lab's inference from observed fan-out queries, not documented mechanism, and a simpler explanation exists: fan-outs are generated from the prompt, so a category prompt naturally produces queries naming well-known category members. Nobody outside OpenAI can settle it. Usefully, the practical advice is identical either way — earn third-party mentions.
You also can't verify your own contribution to the weights; there is no training-data console. The nearest proxies are free-recall probing, of the sort DEJAN used when it ran 200,000 brand-recall surveys against a Gemini model, and reading what an engine says with search off. Both measure the output of memory, not your input to it — which is why measuring both modes continuously beats theorising about causes.
Two playbooks, two clocks
Once you accept the split, the roadmap sorts itself. These are different jobs, with different feedback loops and different failure modes.
| Retrieval half | Memory half | |
|---|---|---|
| What moves it | OAI-SearchBot access, pages written to fan-out sub-questions, natural-language slugs, Google rankings, placement in third-party ranked lists |
Volume and breadth of third-party mentions; presence on Wikipedia, YouTube, Reddit, LinkedIn and review networks; category-level PR |
| Feedback speed | Days to weeks (median 6.81 days to first citation) | Months to a model generation (cutoff currently 16 Feb 2026) |
| Verifiable? | Yes — citation rate, cited URL lists, referral traffic | Only indirectly — mention rate on answers where no search fired |
| Durability | Fragile: a model release can undo it | Slow to build, slow to correct, durable once earned |
| KPI that proves it | Citation rate on retrieval-triggering prompts | Mention rate on prompts where no search fired |
That last row is the operational point. A mention with no citation is a memory-side win; a mention with a citation is a retrieval-side win. Collapse the two into one number and you lose the ability to say which half is broken — which is why Geoptimizer weights mention rate at 35% and citation rate at 25% as separate components rather than folding them together.
The durability row deserves its own warning. When ChatGPT switched to GPT-5.3 Instant on 4 March 2026, RESONEO's tracked prompts saw average unique domains cited per response fall from 19.1 to 15.2, about 20.5%, with the URLs-per-domain ratio holding at 1.26. Same pie, fewer slices — and no site changed anything to cause it.
How to measure without fooling yourself
Four setup decisions determine whether your numbers describe a buyer's experience or an artefact of your own tooling.
Auto, not forced. Benchmark with the search tool forced on and you measure the retrieval half only, producing a rosier, more link-shaped picture than any real user gets. Diagnose with tool_choice: "auto" and record whether a web_search_call actually appeared.
App versus API is not a detail. In the April 2026 observational study, Reddit "received exactly zero AI citations via API" while appearing in 8.9–15.6% of citations in web UI responses from the same platforms. That gap is wide enough to reverse a strategy — imagine launching or killing a community programme on the wrong side of it. Geoptimizer queries the official APIs with web search enabled and states openly that API answers match app answers "closely, but not exactly," which is the kind of disclosure worth demanding from any tool whose numbers you compare. It's also why its Frontier Deep Scan runs on the flagship models the consumer apps ship with — the closest read available to what a buyer actually sees.
Prompts don't live alone. A study of 180 conversations published in August 2026 found full-conversation answers differed materially from isolated final-message answers in 44.7% of cases (95% CI 33.8–56.1%), with 26.7% involving changed recommendations — strongest in the commercial corpus at 68.5%. Your isolated tracked prompt is a clean sample, not a faithful reproduction of a buyer's fifth message.
Repeated runs beat more prompts. A 50-prompt set run 10 times tells you more about stability than a 500-prompt set run once, and citation-share numbers belong in bootstrap intervals — call a trend only when this week's interval doesn't overlap last week's. Engine structure matters too: Semrush's index of 126 million US AI search prompts from January to April 2026 found ChatGPT averages around 15 sources per response against Gemini's three, with brand-mention and cited-domain overlap on Gemini as low as 30%.
Finally, re-baseline around model releases instead of reacting to them. The 7 May 2026 branded-link change shows why: the share of ChatGPT responses containing brand URLs jumped from around 4.5% to 20–24%, and B2B Software and SaaS daily referrals rose more than 200% while ecommerce stayed flat. A visibility metric that barely moved produced a traffic line that tripled. For the full diagnostic, see our walkthrough on telling a genuine visibility shift from a model-version artefact.
What to do Monday
None of this needs a new budget line. It needs one new column and one honest sort.
- Add a "search fired?" column to your tracking sheet. Run each tracked prompt and log the binary. Two runs a week for a fortnight gives a usable per-prompt trigger rate.
- Sort the prompt set by that rate. High-trigger prompts — discovery, comparison, review-seeking, "best X for Y" — go to the retrieval playbook; definitional, category-education and problem-stage prompts go to the memory playbook. If your set skews low-trigger, the set may be the problem: our guide to choosing the 25 prompts worth tracking covers building one that maps to a funnel instead of a keyword export.
- Confirm
OAI-SearchBotcan reach you. Opting it out removes you from ChatGPT search answers entirely, and changes take about 24 hours to register. The free AI Crawler Checker covers GPTBot, ClaudeBot and eight other crawlers in one pass. - Baseline both halves separately. Mention rate where no search fired is your memory read; citation rate where it did is your retrieval read. Track them side by side, never as one average.
- Re-baseline after every frontier release rather than rewriting pages in response to a number a routing change moved.
For the fastest version of step four, Geoptimizer runs your buyer prompts live on ChatGPT, Gemini, Claude and Grok with web search enabled and returns a per-engine breakdown in about 90 seconds, keeping mention rate and citation rate as separate published components rather than one opaque number. Every plan includes all four engines, including the free one — enough to find out which half of your visibility is broken before you commit another quarter of content to it.
FAQ
How often does ChatGPT search the web? About a third of the time, with wide error bars. Semrush's clickstream analysis found search enabled on 34.5% of queries in February 2026, down from 46% in late 2024, with the share ranging from 15% to 66.3% across the study. Profound's conversation-level read puts around 18% of conversations at one or more web citations; Nectiv measured 31% of prompts triggering a search. OpenAI publishes no official figure.
Does ChatGPT search more for commercial and comparison prompts? Substantially more. In an April 2026 study, 391 brand and product queries run through the ChatGPT web UI triggered a search 42% of the time, and the researcher's API dataset showed roughly 73% for discovery, 70% for review-seeking and 65% for comparison, against about 10% for purely informational queries. The population average is dragged down by the 37.5% of prompts that are generative "write me X" tasks.
If ChatGPT doesn't search, can I influence the answer at all? Not this quarter, and not with a page. Answers from parametric memory reflect the training corpus, and GPT-5.6's stated cutoff is 16 February 2026. The levers are third-party mentions at volume: Victorious found brands with fewer than 2,000 indexed pages mentioning them were named just 3% of the time, and Ahrefs measured branded web and YouTube mentions correlating with AI visibility far more strongly than backlinks.
Why do my new pages get crawled but never cited? Crawl latency is rarely the cause — median time to first ChatGPT or Claude citation was 6.81 days, with 90% inside about 37 days. Format and ownership explain more: 63% of AI citations across engines go to listicles, rising to 70.8% for "best"/"top" pages in B2B SaaS prompts, and 99.99% of citations analysed in category-research prompts pointed to third-party domains.
Should I force web search on when benchmarking AI visibility?
Only if you specifically want to measure the retrieval half. Forcing a search with tool_choice: "required" guarantees an outcome real users get part of the time, inflating how link-shaped your visibility looks. Run diagnostics on auto and log whether a search fired.
The short version
About a third of ChatGPT queries touch the live web, and far more than that on the commercial prompts your buyers actually type — but the answers that never search still recommend somebody, and nothing you publish this quarter can reach them. Retrieval is fixable in days and fragile to model releases; memory is fixable in quarters and durable once earned. Teams get stuck by spending retrieval-shaped effort on a memory-shaped gap, and they stay stuck because their dashboard reports one number instead of two.
Sort your prompts by whether a search fires, split the roadmap accordingly, and measure mention rate and citation rate as separate outcomes. Run a free AI visibility check across all four engines and find out which half of yours needs the work.