· Updated · 19 min read · Geoptimizer Team

How to Choose the 25 Prompts You Track for AI Visibility

  • generative-engine-optimization
  • ai-visibility
  • prompt-tracking
  • measurement
  • geo-strategy
How to Choose the 25 Prompts You Track for AI Visibility

Choose 25 distinct buyer intents mapped across your funnel — not 25 phrasings of the same question, and not your keyword list rewritten with question marks. The prompt set is the measurement instrument, so its composition decides your score before any engine answers: in one study of 1,248 commercial prompts across 52 B2B SaaS categories, challenger brands held just 14% share on generic "best [category]" prompts but 41% on startup and mid-market prompts (MaxAEO). Same brands, same week, same engines — a 27-point swing produced entirely by prompt selection.

That is the uncomfortable part of setting up your first generative engine optimization tracker. You are not reading a ranking; you are drawing a sample, and you drew it yourself. This guide covers what 25 slots actually buy you statistically, why your keyword list is the wrong seed, an explicit funnel allocation for all 25, where to source buyer wording, and the five checks a candidate should pass before it takes a slot.

TL;DR:

  • Spend all 25 slots on distinct buyer intents — phrasing duplicates waste the budget, and repetition should come from refresh cadence instead.
  • Prompt selection alone can swing measured share 27 points, and "best X" phrasings inflate mentions by roughly 20% — rewrite generated drafts with real buyer situations and constraints.
  • Use the explicit allocation: 5 category-head, 7 use-case, 4 persona, 4 comparison, 3 problem-stage, 2 brand — about 75/25 unbranded to branded.
  • Source wording from real buyer language: sales calls, forums, Bing Webmaster and Search Console question queries.
  • Freeze the set once it's live, review quarterly, and annotate every change — a changed prompt set breaks the trend line it was measuring.

Why most first prompt sets are wrong

The default path is to open a GEO tool, accept its suggested prompts, and start measuring. Those suggestions are a reasonable draft. They are also measurably different from what humans type.

Otterly compared hundreds of real prompts — collected through a US market survey of SEO and marketing professionals — against prompts estimated by prompt-research tooling and Search Console data. Real prompts averaged 15.1 words (median 12); estimated prompts averaged 8.8 (median 7). Real prompts contained personal pronouns 52.1% of the time versus 18.8%, and were conversational in 23.9% of cases versus 5.8%. The tell is in the first word: estimated prompts most often start with "best" (18.8%), while real prompts start with "what" (15.5%) or "I" (11.3%) (Otterly).

That gap is a category problem, not a tooling defect. "Best CRM software" is a keyword. "I run ops for a 12-person agency and our CRM keeps getting abandoned after trial — what should we look at?" is a prompt. The second contains a situation, a constraint and a pain point; the first contains a category. Peec's own documentation names this as the most common first-tracker mistake — teams "treat AI prompts like traditional SEO keywords" (Peec docs).

There is a second, subtler failure: format bias. In a study of 1,754 prompts and 37,804 AI responses across ChatGPT, Gemini, Perplexity, Google AI Mode and AI Overviews, ranking or list-format prompts produced about 20% more brand mentions than open-ended questions (Search Engine Journal on Peec's research). A set stuffed with "best X" prompts reads higher than a set of natural questions — instrument bias, not visibility, and one of the structural reasons two GEO trackers report different scores for the same brand.

As Cate Dombrowski of Omniscient Digital puts it: "The results you get from LLM prompt tracking are only as valuable as the prompts you choose as input" (Omniscient Digital).

What 25 prompt slots actually buy you

Before allocating slots, know their statistical worth. Most setup guides skip this, and it changes every downstream decision.

Visibility — the proportion of answers in which your brand appears — behaves like a probability, so it carries a margin of error. The standard formula is MoE = 1.96 × √(p̂(1−p̂)/n). At a 30% mention rate, that gives ±16.4pp at n=30, ±9.0pp at n=100, ±4.5pp at n=400 and ±2.8pp at n=1,000 (MaxAEO).

Now the arithmetic for a 25-prompt plan. Because every active prompt runs on all four engines on Geoptimizer's Pro plan, 25 prompts is 100 answers per full refresh — not 25. At p ≈ 0.30 that single refresh gives roughly ±18pp per engine (n=25) and ±9pp pooled (n=100). Every prompt is re-checked weekly, so ~4 scheduled refreshes a month produce ~100 answers per engine — naively ±9pp, but repeated runs of one prompt are correlated. Applying a design effect of 1 + (m−1)·ICC at a mid-range ICC of 0.5, the effective sample falls to roughly 40, or about ±14pp per engine per month from scheduled coverage alone; on-demand spot-checks and full sweeps tighten it further. (Formula and the ICC range of 0.44 on Gemini to 0.73 on Perplexity are MaxAEO's; the plan-specific arithmetic is ours.) The same margin-of-error machinery, applied to the gap between you and a competitor rather than your own rate, is worked through in Competitive AI Visibility Analysis.

A 25-prompt set is therefore a monitoring instrument: it reliably catches moves of 10pp or more and tells you presence or absence on each specific prompt. It cannot resolve a three-point change — which is why a headline score should be a rolling window with a confidence band rather than a single scan.

Breadth beats depth — at a fixed budget

The natural instinct is to hedge by tracking three phrasings of your most important question. The math argues against it. At a fixed budget of 600 answers, MaxAEO found 20 prompts × 30 runs yields ±15.4pp while 200 prompts × 3 runs yields ±5.4pp — a 3.3× spread in precision from allocation alone, because repeated runs of one prompt are correlated and add less information than fresh prompts do.

The counterweight is real. Nick Lafferty argues the opposite: "Repeated runs beat more prompts: a 50-prompt set run 10 times tells you more about stability than a 500-prompt set run once" (AI visibility metrics reference), and Evertune found that repeating one prompt five times carries a 27-point margin of error, dropping to 12 points at 12 repetitions (Evertune). They are answering different questions: MaxAEO estimates category-level share, Lafferty asks whether one prompt's result is real. For a 25-slot plan the resolution is clean:

Spend all 25 slots on distinct intents. Let repetition come from refresh cadence over time, never from duplicate prompts.

Breadth on day one, depth by week four, without spending a slot twice. The one exception: phrasing variation matters most for unbranded, commercial mid-funnel queries, while top- and bottom-of-funnel prompts are comparatively stable (SEJ/Peec). If you ever do spend a second slot on a variant, spend it there.

Why your keyword list is the wrong seed

Most first prompt sets are built by exporting the top 25 keywords and adding question marks. It is the fastest path and it produces a set that measures the wrong market.

EMGI studied 150 SaaS companies across 120 keywords in six categories (data collected 5–8 April 2026) and found 44% of SaaS brands ranking in Google's top 10 received zero ChatGPT mentions for the same keywords, while 81% of the brands ChatGPT recommended were not in Google's top 10 (EMGI). Category gaps varied widely — 53% in marketing automation and 52% in analytics, but only 18% in dev tools.

The counterweight in the same study keeps this honest: keyword-ranking count still correlated with AI citations at r = 0.76, far above organic traffic (r = 0.23). Your keyword list is a legitimate starting point for topic selection and a poor source of wording. As a vendor-adjacent study, treat the direction as reliable and the exact percentages as indicative.

Two more mechanical differences separate keywords from prompts:

Intent mix. Profound's classification of over 50 million ChatGPT prompts found commercial intent at 9.5% and transactional at 6.1%, versus 14.5% and 0.6% in traditional search — transactional intent runs roughly an order of magnitude above its share of classic search (Profound). The commercial slice is smaller but far more decision-shaped.

One prompt is not one query. Peec analysed over 20 million ChatGPT query fan-outs between October 2025 and January 2026 and found 2.3–2.8 machine-written queries per prompt across five countries (Peec). Optimising a page for a tracked prompt's literal wording means optimising for a string the engine rewrote before it searched.

This is also why prompt tracking exists at all. Similarweb's US desktop panel (July–December 2025) found users were 2.5× more likely to visit an AI-recommended brand than its direct competitor within seven days — yet 55.9% of AI-influenced visits appeared in analytics as search traffic (Search Engine Land). Big effect, mostly invisible referral trail, so the prompt set is your leading indicator.

To translate that AI-driven lift into measurable site outcomes, follow this walkthrough on how to connect AI visibility to GA4 that maps visibility scores to downstream traffic and conversions.

The funnel map: how to allocate all 25 slots

Here is the allocation, deliberately opinionated so you can adapt it rather than start from a blank page. The logic is one slot per distinct intent, weighted toward the constrained mid-funnel prompts where challengers actually appear.

Bucket Slots What it looks like Why this many
Category head 5 "Best [category] software 2026", "top [category] tools for [industry]" High volume, high difficulty. Your baseline against incumbents — but challenger share here is only 14%
Use case & constraint 7 "[Category] tool that handles [specific workflow] with [constraint]" Challenger share rises to 34–38% on use-case, integration and compliance prompts
Persona & segment 4 "…for a 15-person startup", "…for a mid-market team on a $500/mo budget" Highest challenger share in the data at 41%
Comparison & alternatives 4 "[Competitor] alternatives", "[A] vs [B] for [use case]", "migrating off [incumbent]" Migration prompts show 37% challenger share; closest to a decision
Problem stage 3 "Why does our team keep abandoning [tool type] after trial?" Catches buyers before they name a category
Brand & accuracy 2 "What is [your brand]?", "Is [your brand]'s pricing…?" Verifies engines describe you correctly

Total: 25. Comparison and brand prompts together are 6 of 25 — about 24%, landing on Conductor's recommended ratio of roughly 75% unbranded / 25% branded. Conductor names "over-relying on branded prompts that inflate visibility metrics" as a common mistake (Conductor Academy), and there is hard evidence for the inflation: in Otterly's own Bing grounding-query data, branded queries carried roughly twice the citation density of non-branded ones (75 versus 35 citations per query).

The allocation cross-checks against published sets. SE Ranking recommends 20–40 prompts total, split into 10–20 awareness, 20–30 consideration and 5–10 brand-evaluation prompts tracked separately (SE Ranking). Kevin Indig's production structure is 40 seed prompts — 12 brand, 12 category, 16 problem-focused — with three personas applied to the 28 category and problem prompts (Growth Memo); at 25 slots persona becomes a qualifier baked into individual prompts rather than a separate dimension. On length, Radyant recommends 70–80% long, context-rich prompts of roughly 180–200 characters (Radyant) — consistent with Otterly's 15-word real-prompt average.

One refinement worth stealing from Omniscient Digital's 200-prompt research set: control qualifier density by stage. They kept under 10% of prompts qualified at problem-unaware, about 30% at problem-aware and roughly 60% at solution-aware. So your three problem-stage prompts should be nearly bare, while persona and use-case prompts should be thick with constraints.

Where to get wording that sounds like a buyer

Allocation gives you 25 empty slots. Filling them with buyer language is the part that separates a useful set from a keyword list in question clothing.

Bing Webmaster Tools grounding queries. The AI Performance report — public preview 9 February 2026, expanded 16 June 2026 with intent labels, topic groups and citation share — exposes the reformulated queries Copilot generates behind the scenes, plus per-page citation counts. It covers Microsoft Copilot only, and those are machine queries rather than user prompts, but it is free observed retrieval data (Otterly's walkthrough).

Google Search Console — with a clear limit. Google's generative-AI performance report, announced 3 June 2026 and rolling out first to a subset of UK site owners, reports impressions, pages, countries, devices and dates for AI Overviews and AI Mode. It gives no clicks and no query data (Search Engine Land). Your existing GSC question queries — filtered by a regex on why/what/when/how/should/can/does — remain the better source there.

For a practical walkthrough on splitting AI visibility by country and language and choosing market‑specific prompts, see AI Visibility by Country and Language: Measure Per Market.

For a tactical breakdown of which signals matter and how Gemini, AI Overviews and AI Mode differ in what they cite and prioritize, see Gemini vs AI Overviews vs AI Mode: What to Track.

Forums and communities. Ahrefs' guide lists eleven sourcing methods, including forum discussions via Google's &udm=18 parameter, Perplexity related questions, People Also Ask, pages already receiving AI-bot traffic in your logs, and persona queries built on the template [Situation][Constraints][Priorities][Pain points][Question] (Ahrefs). That template is the highest-leverage artefact here — it produces exactly the pronoun-heavy, 15-word prompts Otterly found in the wild.

Your sales calls. Radyant's framing is blunt and correct: "30 minutes with the right person tells you more about your audience than weeks of keyword research." Listen to three recorded discovery calls and note how buyers describe the problem before they know your category exists. That is where your problem-stage prompts come from.

Vendor prompt databases — read the fine print. Ahrefs' Brand Radar advertises "150+ million monthly potential prompts and AI responses" across six AI indexes, powered by "actual questions and PAAs from our 110 billion keyword database" (Ahrefs Academy). That is search-derived question data, not chat logs — useful, but not observed prompt demand.

Generated prompts as a first draft. Geoptimizer's free Prompt Ideas Generator turns your domain into 20 buyer-intent prompts, no signup required. Given the Otterly data, the workflow is: generate, then rewrite. Add the pronoun, the team size, the budget, the incumbent you are migrating from. Generation gets you to the intent in seconds; your rewrite gets it into buyer language.

Five checks before a prompt takes a slot

Run each candidate manually once before you pin it. Five minutes per prompt is cheap insurance on a set you intend to freeze for months.

1. Does it trigger retrieval? Nectiv analysed over 8,500 ChatGPT prompts across nine industries and found only 31% triggered at least one web search (Search Engine Land). Commercial intent fares far better — Profound found commercial prompts trigger web search 53.5% of the time versus 18.7% for informational. A prompt that never triggers retrieval cannot produce a citation, so it is a half-wasted slot if citation rate is part of your score.

2. Does any brand get named? Peec's docs note that informational prompts rarely produce brand names unless brand context is added. Run it; if the answer names nobody, the prompt measures nothing.

3. Could you realistically appear? SE Ranking calls this competitive relevance. If the answer lists five enterprise incumbents and you are a 20-person company, that prompt will read zero for months. Keep one or two as aspirational benchmarks; do not build a set out of them.

4. Is it influenceable? Look at which sources the engine cites today and ask whether you can plausibly earn placement there. This varies sharply by engine — Ahrefs tracked over 1.9 million US queries in June 2026 and found that on Grok, Reddit was the most-cited domain at 16.3% of mention share, with YouTube (15.1%) and Facebook (13.9%) taking the top three to 45.3% of citations (Ahrefs).

For a detailed breakdown of how Grok selects and cites sources — and what that implies for your ability to influence its answers — see How Grok Chooses Its Sources: The X Citation Channel.

For practical tactics on earning placement on the specific "best tools" lists that engines tend to cite, see our guide on how to get into the best tools lists AI engines cite.

5. Does it map to a revenue motion? Conductor's rule is the cleanest heuristic here: "If you wouldn't invest content resources into a topic, don't track it." A prompt you would never write a page for is a prompt whose score you cannot move.

Rules that keep the number honest over time

The set is built. These five rules protect the trend line.

Freeze the set. "Freeze the prompt set for comparable periods; changing prompts changes the denominator" (Lafferty). A 25-slot plan tempts constant swapping — resist it. Review quarterly, change a few slots at most, and annotate the date of every change. Geoptimizer counts only active prompts against your limit and frees a slot the moment you pause one, so you can retire a prompt without deleting its history — the same archive-don't-delete practice Peec's docs recommend.

Report per engine, always. Semrush's 2026 AI Visibility Index found ChatGPT averages about 15 sources per response while Gemini averages about 3 (Semrush) — a five-fold difference in citation surface for the same prompt. A prompt that looks dead in a combined score may be healthy on ChatGPT and simply a poor fit for Gemini's narrower sourcing. The per-engine breakdown is how you tell a bad prompt from a bad engine fit.

Keep branded and unbranded in separate numbers. Mixing them is on SE Ranking's mistakes list, and the 2× citation-density gap in Otterly's grounding data shows how much a few branded prompts can lift a blended average.

Judge moves against a band, not a scan. SparkToro's research — 600 volunteers, 2,961 prompt runs on ChatGPT, Claude and Google AI — found under a 1-in-100 chance that two responses return the same list of brands, and roughly 1 in 1,000 for the same order (SparkToro). Rand Fishkin's conclusion: "Any tool that gives a 'ranking position in AI' is full of baloney." Geoptimizer reports a 7-day rolling window with a confidence band for precisely this reason, and labels single on-demand scans as snapshots. If two windows' bands overlap, you have no measurable trend yet — the first thing to check before concluding a new frontier model release moved your visibility.

Expect the underlying sources to churn. Only 2.3% of citations persist after three identical prompt runs, Google AI Mode replaces about 56% of cited sources weekly and ChatGPT about 74% (Growth Memo, 8 June 2026). Kevin Indig's framing is the right mental model: "The next iteration of prompt tracking will look less like rank tracking and more like polling: repeated runs, clear sampling rules, confidence intervals, segmented panels, and raw-answer audits."

Calibration: what a good number looks like at 25 prompts

DerivateX's 2026 B2B SaaS benchmark ran 7 prompts each for 50 companies across four platforms (1,400 prompts, March–April 2026) and found an average AI Presence Score of 56.9/100, median 63.5; 44% scored below 50 and 78% appeared on all four platforms (DerivateX). On timing, Profound's data puts median time-to-first-citation for a new page at 6.81 days, with a P90 of 37.1 days.

Two limitations belong on the record. Trackers query engines in a clean, memory-free state, but since OpenAI's 5 May 2026 memory-sources update ChatGPT can draw on past chats, saved memories and — for Plus and Pro users — connected files and email, so two identical users can receive different recommendations. A tracker measures the baseline every user starts from, the only comparable control available. And several of the best numbers here come from tool vendors and agencies; they are usable because they publish methodology and sample sizes, which is the standard Fishkin asks buyers to demand.

When do you outgrow 25 slots? The trigger is not "I want a bigger number." It is: your list of genuinely distinct buyer intents is longer than 25. If you serve three segments with different constraints, or track migration prompts against four incumbents, you will hit that wall quickly, and the Pro plan — a 30-prompt tracked pool, each prompt re-checked weekly — buys both more intents and a tighter band. Until then, 25 distinct, well-sourced, frozen intents beat 100 rushed ones — and more measurement fundamentals are covered across the blog.

FAQ

How many prompts do I need to track AI visibility?

For a first tracker, 20–40 distinct prompts is the practitioner consensus (SE Ranking recommends 20–40; Kevin Indig runs 40). What matters more is answers, not prompts: 25 prompts across four engines is 100 answers per refresh, roughly ±9pp pooled margin of error on a single scan at a 30% mention rate. Give it 30 days before drawing conclusions.

Should I track the same prompt with different wording?

Usually no. Peec's analysis of 1,754 prompts found 88–92% of human phrasing variants for the same commercial intent sat above 0.50 cosine similarity, with brand mention frequency staying consistent above that band — so three phrasings of one question buy you very little. The exception is unbranded mid-funnel commercial queries, where phrasing sensitivity is highest.

Can I just convert my Google Search Console keywords into prompts?

Use them for topic selection, not wording. EMGI found 44% of SaaS brands in Google's top 10 got zero ChatGPT mentions for the same keywords, and 81% of ChatGPT's picks were not in the top 10 — although keyword-ranking count still correlated with AI citations at r = 0.76. Take the topics from GSC, then rewrite with the situation, constraint and pain point a buyer would include.

How often should I change my prompt set?

Review quarterly and change a few slots at a time. Changing prompts changes the denominator, so every swap creates a break in your trend line. Annotate the date of each change, pause rather than delete so history survives, and never edit the set in the middle of investigating a score move.

How many branded prompts should be in the set?

Roughly a quarter of the set at most, and reported as a separate number. Branded queries carried about twice the citation density of non-branded ones in Otterly's grounding data, so a branded-heavy set produces a flattering average that tells you nothing about whether new buyers can find you.

Start with 25 intents, not 25 keywords

The method compresses to five decisions: allocate slots by funnel stage rather than by search volume, spend every slot on a distinct intent, source wording from buyers instead of keyword exports, qualify each candidate against the five checks, then freeze the set and judge moves against a confidence band.

If you want to see what your current set produces across all four engines before committing to it, run a scan — Geoptimizer includes ChatGPT, Gemini, Claude and Grok on every plan including the free tier, publishes the full scoring formula as a versioned composite with a changelog, and returns a first score in about 30 seconds. Start free and check your prompt set against real answers — then upgrade only when your list of distinct intents outgrows your slots.

Keep reading

See it on your own domain.

Free visibility check across ChatGPT, Gemini, Claude, and Grok — about 30 seconds.

Run the free check