· Updated · 20 min read · Geoptimizer Team
Competitive AI Visibility Analysis: Share of Answers
- generative-engine-optimization
- ai-visibility
- measurement
- competitive-analysis
- share-of-answers
A competitive AI visibility analysis measures what percentage of your category's prompts each rival owns inside AI answers — not how often the category gets talked about. The distinction is expensive to ignore: when Lily Ray analysed 100 B2B "best [category] software" queries in Google AI Overviews, self-promotional listicles were cited 323 times across 80 triggered AI Overviews, but in 224 of those instances Google cited a brand's own page without recommending that brand — competitors captured the recommendation 69% of the time. Being in the sources and being the answer are two different competitive events, and only one of them wins the deal.
The clearest example in that study: for the query "best LMS for selling courses," Google cited Oasis LMS's own article and then recommended Kajabi, Thinkific, LearnWorlds and Teachable — all brands named inside the Oasis article. A share-of-voice report would have logged that as a win.
This is a method, not a dashboard tour — if you're still sorting out what generative engine optimization even is, start there. Five steps, each one designed to survive the meeting where your VP asks how you know the number is real.
Why share of voice doesn't survive the trip to AI answers
Share of voice was built on a countable universe. You knew the keyword set, you knew roughly how many searches happened, and everyone measuring the category was measuring approximately the same thing.
That premise is gone. As Dan Taylor put it in Search Engine Land in June 2026, "the universe of possible AI prompts is effectively infinite" — which means every vendor percentage you have ever seen is a share of answers within someone's chosen prompt set. Taylor's critique is that these arrive as "precise-looking percentages that can't be audited or validated," and his fix is to split the single number into share of mentions, share of recommendations and share of narrative.
Here is the reframe, and it is the load-bearing idea of the whole method: share of answers is only defensible as a relative, within-set metric. "We are at 22%" is unfalsifiable. "We appear in 22% of our 120 frozen category prompts; Competitor B appears in 41%" is a real finding, because the denominator is identical for both of you and published. Every serious critique of AI share of voice lands on the absolute number and the hidden denominator — none of them damage a comparative reading against a frozen set. The objection becomes the reason for the method.
The second thing that breaks in transit is the assumption that a mention and a citation are the same event. In the Semrush and Kevin Indig ghost-citations study — 3,981 domain appearances across 115 prompts and 14 countries in June 2026 — 61.7% of citations produced no brand mention at all. Only 13.2% of brand appearances included both a citation and a mention; 25.1% were mentions with no citation. The platform used the page and never named the brand. As Indig described it, "the AI knows the information came from somewhere, but doesn't feel the need to explicitly say so to users."
That split is not evenly distributed. In the same study ChatGPT ran 87% citation-leaning against a 20.7% mention rate, while Gemini was almost the mirror image at 21.4% citations and 83.7% mentions. A scoreboard that collapses the two into one figure will tell you that you and your rival are level when one of you is being read and the other is being recommended. So: two numbers per brand per engine, always.
Step 1 — Freeze the competitor set and the prompt set
Do this before you look at a single data point, because both choices are trivially easy to rationalise after the fact.
The competitor set. The pattern that shows up repeatedly in published competitive-profiling work is 3–5 direct competitors plus 2 aspirational leaders, run against 50–100 category prompts on a monthly cadence. Two rules matter more than the count:
- Segment out general platforms. Amazon, Reddit, YouTube and Wikipedia will appear constantly in your results. They are source domains, not rivals. Leaving them in the competitor column inflates the denominator and makes your gap look structural when it is just the shape of the web.
- Check for name collisions before you trust anything. Practitioner reporting describes a startup whose 30 tracked AI mentions turned out to be 28 references to a fictional character sharing its name — its "top performing topics" were fan theories and costume guides. Even the academic literature gets caught: the authors of the June 2026 arXiv study below had to void EY's 95.5% category ownership score in consulting as an artefact of substring contamination in alias matching. Run every brand string — yours and every rival's — against a sample of raw answers before you compute anything.
The prompt set. This is the methodology. A high share of answers can be manufactured by loading the set with easy definitional prompts, so the construction rules need to be written down:
- Start at roughly 100 prompts — Profound's published guidance, with users typically running 100–1,000 tracked prompts built from SEO inputs plus sales and support language.
- Cover awareness, evaluation, conversion and post-purchase. A set weighted to one stage measures one stage.
- Track far more unbranded than branded prompts. A prompt containing your own brand name measures prompted recall, not organic share. Keep head-to-head comparison prompts in a separate bucket so they cannot inflate the headline.
- Version the set and freeze it between reporting periods. Comparing this month's number to last month's different set is not a trend, and it is the first thing a sceptical finance partner will find. The board-ready form is a locked set of 40–80 buying-stage questions versioned for review.
Phrasing sensitivity is real and you will get asked about it. In testing by Unusual, a single brand's share of voice shifted by as much as 17% when single words were swapped for synonyms, and only 16 of 100 prompt pairs produced identical vendor results. The answer is not to abandon the metric but to build prompt clusters: three or four phrasings of the same buying intent, averaged together, so the cluster is the unit of analysis rather than any single string.
Step 2 — Choose an allocation, then attach a margin of error
Most competitive AI visibility reports fail here. They present a gap without a confidence interval, and the first person who notices the score bounced 6 points last week stops believing all of it.
The rate you are measuring is a proportion, so the standard formula applies:
MoE = 1.96 × √(p̂(1 − p̂) / n)
At a 30% mention rate, that gives ±9.0pp at n=100 answers, ±6.4pp at n=200, ±4.5pp at n=400 and ±2.8pp at n=1,000.
The non-obvious part is that how you spend a fixed budget matters more than the size of the budget. Repeated answers to the same prompt are correlated, so they do not each buy a full unit of precision. The correction is the design effect, Deff = 1 + (m − 1) × ICC, where m is runs per prompt and the median within-prompt intra-cluster correlation across engines is 0.57 — from 0.44 on Gemini up to 0.73 on Perplexity.
Apply that to a fixed budget of 600 answers and the spread is dramatic:
| Design | Effective n | 95% MoE |
|---|---|---|
| 20 prompts × 30 runs | 34 | ±15.4pp |
| 100 prompts × 6 runs | 156 | ±7.2pp |
| 200 prompts × 3 runs | 280 | ±5.4pp |
Same cost. Roughly a 3× difference in precision, from allocation alone.
Reconciling the run-count argument you will be shown
There is a live public disagreement about run counts, and someone will forward it to you. Both sides are right about different things.
The depth-first case: Evertune, publishing on 16 July 2026 from 10,700 prompts on ChatGPT, reports that a single prompt carries a margin of error of roughly ±27 points at 5 repetitions, ±12 at 12, and ±6 at 100, with diminishing returns beyond that. Their framing: "Ask one the same question twice, and you will get two different answers."
The breadth-first case: Profound ran 753 prompts across 7 US platforms from 1–14 June 2026, once daily (~129,000 runs) versus ten times daily (~860,000 runs). Result: visibility of 78.7% versus 80.4%, and citation share of 10.24% versus 9.99%. The typical day-to-day difference was about 2 percentage points for visibility and 0.3 points for citation share — a 10× cost increase for a rounding error, against a drift floor caused by the platforms themselves changing.
These are not in conflict, and noticing why is the most useful thing you can bring to your next planning meeting. ±27pp is the error on a single prompt's rate. 1.7pp is the error on a 753-prompt portfolio rate. Pooling across many prompts is itself a form of repetition. Both are correct; they describe different quantities.
The practical rule that falls out:
- Breadth for the headline number you show leadership. Roughly 150–200 prompts at 3 runs, weekly, lands near ±5pp.
- Depth for the 15–25 prompts you are actively trying to win. Nine to twelve runs each, because at prompt level a single reading is close to worthless — in a 23,040-answer study across 18 B2B SaaS brands, 320 prompts and 6 engines, only 34% of prompts returned identical brand shortlists across all runs, 41% flipped the tracked brand in or out at least once, and single runs disagreed with the 12-run majority 19% of the time.
For practical guidance on picking those 15–25 priority prompts and balancing runs versus coverage, see How to Choose the 25 Prompts You Track for AI Visibility.
This maps onto how scan types work in practice. In Geoptimizer, a full sweep re-runs every active prompt across all four engines — the breadth pass behind the headline number — while a spot-check re-runs only the prompts you have pinned as priorities, the depth pass on the slots you are contesting. The headline score is reported as a 7-day rolling window with a confidence band rather than a bare point estimate, which is the same instinct the math above demands. For the underlying reasoning on why two honest tools can disagree about the same brand, our breakdown of why AI visibility tools return different scores covers the five legitimate sources of divergence.
Step 3 — Compute share of answers per engine and per cluster
Resist the blended number until the very last step. Three findings make engine-level reporting the minimum viable method rather than a nice-to-have.
First, engines disagree about who leads. In a study of 250 brand-neutral category queries run five times each across GPT-5.2, Gemini 3 Flash and Perplexity sonar-pro — 3,750 responses covering 50 brands in 5 industries — all three models picked the same top brand in only 41.6% of queries, ranging from 22.0% in consulting to 73.8% in e-commerce.
If you need a practical checklist for engine-specific signals and which Gemini surfaces to prioritise, see what to track across Gemini, AI Overviews and AI Mode.
Second, they draw on different source pools. Evertune's review of roughly 25,000 unique URLs and nearly 400 million citations across six engines found Gemini-powered surfaces overlapping more than 50% with each other, while Copilot overlapped just 4–6% with other models.
Third, the number of slots per answer differs mechanically: ChatGPT averages around 15 sources per response while Gemini averages around 3, per Semrush's 2026 AI Visibility Index. A rival winning 40% of ChatGPT's fifteen slots is a different problem from one owning two of Gemini's three.
Because ChatGPT surfaces far more slots, it's important to measure paid placements alongside organic mentions; for a practical guide to measuring ChatGPT ads and organic AI visibility, see measuring ChatGPT ads and organic AI visibility.
A ready-made scoreboard
You do not need to invent metrics. The June 2026 arXiv paper supplies three that a challenger brand can copy directly:
- Category Ownership Index —
COI(b,q) = mention_count(b,q) / total_iterations(q). Brand b's share of mentions for prompt q. Roll it up to cluster, then category. - Competitive Vacuum Index —
CVI(q) = 1 − max COI(b,q). Values above 0.50 mean no brand dominates that prompt. This is your target list. - Displacement Score —
DS(A,B,q) = P(A|¬B) − P(A|B). Positive means you appear more often when a rival is absent (you substitute for each other); negative means you get co-recommended.
That last one is the most under-used competitive metric available, and it is free to compute from data you already collect. The paper found a mean displacement ratio of 2.4:1 across five industries, ranging from 0.4:1 in consulting to 4.3:1 in e-commerce. If your category co-recommends, the job is getting onto the shortlist and the rival is not really the obstacle. If it displaces, the job is knocking someone off it. Same budget, opposite plans.
And for anyone who has been told AI recommendation is winner-take-all: the same paper reports a mean Gini coefficient of 0.28 (95% CI [0.16, 0.41]) and concludes the results "resist simple winner-take-all narratives within the examined scope," with concentration ranging from 0.18 in SaaS to 0.53 in e-commerce.
Step 4 — Diagnose the slot, not just the score
A gap tells you that you are losing. It does not tell you where the slot came from. Five gap types, from the published taxonomy, turn a score into a work queue:
- Mention gap — you are absent from answers where rivals appear.
- Prompt gap — specific questions you consistently lose.
- Source gap — domains that cite competitors but never you.
- Citation gap — you are mentioned, but your site is not the source.
- Narrative gap — you appear, but are described less favourably.
The source gap converts into action fastest. The Semrush walkthrough uses a project-management example: if ChatGPT recommends Asana on the strength of a tech blog's citation, that blog is a source gap for Monday. Filter for domains citing competitors but not you, then prioritise by how often each appears across your prompt set.
This is a finite job rather than an infinite one because AI answers are assembled from a surprisingly small pool. Across those ~400 million citations, 63% pointed to listicles, and 71–86% of those were ranked lists. Company homepages barely register: an analysis of 60,350 citations across 10,000 commercial prompts in 50 industries found only 7% of ChatGPT citations and 4% of Claude citations pointed to a brand's own homepage, with ranked comparison and category pages accounting for roughly 81% and 75% respectively.
Beware: raw crawl counts don't equal AI citation weight — see our analysis of crawler logs and why crawl volume isn't citations.
So the competitive question reduces to something tractable: which 10–20 URLs supply the slots in my category, and where do I rank inside each one?
For a tactical playbook on earning placement inside those ranked listicles and vendor roundups, see how to get into the best tools lists AI engines cite.
Rank inside those lists is measurable and it moves. Peec AI's analysis of nearly 200,000 AI responses across 8 engines between September 2025 and March 2026 found that holding rank #1 in frequently-cited listicles was associated with +16.5pp visibility in B2B SaaS and +13.4pp in emerging MarTech, and with the brand appearing roughly one answer position earlier. These are associations from observational data, not a randomised intervention — but Peec's summary is the right prioritisation heuristic: "Five strong placements in frequently AI-cited third-party sources will matter more than 50 placements in articles AI engines never retrieve."
Two refinements worth building in:
Stage matters. Seer Interactive's analysis of 804,491 AI responses across 1,926 brands and 15,783 prompts found review platforms made up 1.51% of citations at the awareness stage but 24.27% at the intent stage, with Trustpilot the most-cited review site on ChatGPT, Perplexity and Google AI Mode. Brands with fully optimised profiles saw 9.5× more co-mentions than brands with none — though better-resourced brands are also the ones with managed review profiles, so treat it as a signal rather than a guaranteed lever.
Source diagnoses have a shelf life. Semrush tracked 230,000+ prompts and 100M+ citations in weekly snapshots and watched ChatGPT's Reddit citations fall from around 60% of prompt responses to about 10% in roughly six weeks, with Wikipedia dropping from ~55% to under 20% — while Google AI Mode's Reddit citations barely moved. Re-run the diagnosis quarterly at minimum. If the source mix shifts right after a model release, our guide to telling a real visibility shift from model-version noise covers the checks to run before you rewrite anything.
Step 5 — Set a target leadership will actually accept
"Get to 40%" is not a target if the category leader in your sector sits at 16%.
Similarweb's January 2026 analysis of 25,000+ prompts across ChatGPT, Gemini, Copilot and Perplexity in the US gives sector-level anchors for category leaders: Apple at 54.38% in consumer electronics, CeraVe 27.17% in beauty, Reuters 22.98% in news and media, Expedia 18.18% in travel, Nike 16.02% in fashion and Chase 15.89% in finance. The "strong competitive" bands sit far lower — 7–13% in consumer electronics, 15–23% in beauty, 10–16% in travel, 7–10% in finance — with the floor for meaningful presence around 4–8% depending on sector.
Concentration varies enormously, so benchmark against your category rather than a global average. Semrush's index of 126 million prompts found the top 3 brands holding 82.9% of visibility in news and media and 76.9% in consumer electronics, but only 41.4% in finance and 42.2% in industrial. A 12% share means something very different in each.
Then apply a stability rule so you know which gaps are worth chasing. Across 1,094 US categories tracked monthly from January to June 2026, clear owners held position in 90.4% of month-over-month comparisons — but where leadership changed hands, the median lead was 1.3 percentage points, versus 2.9 points where it held. Practical read: a sub-1.5pp gap is contestable this quarter; a ~3pp gap is a moat and needs a different plan.
Two framings that travel well upward. First, excess share of voice: report your share minus your market share, so "we're at 9%" becomes "we're 4 points below our market share." Binet and Field's IPA work associates roughly every 10 points of eSOV with about 0.5% annual market share growth — an imperfect transplant to AI answers, but a frame executives already understand. Second, the challenger's actual opening: only 15.2% of those 1,094 categories had a clear owner, 31.2% had an emerging leader, and 53.7% were unsettled, with no brand appearing consistently. Traditional SEO strength barely predicts who wins — owners beat runners-up on branded search volume just 55.7% of the time, on Authority Score 52.5%, and on organic traffic 48.4%. Roughly a coin flip.
Your unsettled clusters — the high-CVI ones from Step 3 — are a target list, not a mood.
What leadership will challenge, and how to answer it
"AI answers are personalised, so any share figure is fiction." Personalisation and phrasing sensitivity are real. They argue for prompt clusters and repeated runs, not for abandoning measurement — a relative comparison inside a frozen set does not require an accurate estimate of the true prompt universe.
"Prompt volume data is made up." Largely fair. Conductor's July 2026 analysis argues prompt-volume estimates rest on panels representing under 1% of roughly 2.5 billion daily prompts — by some estimates as low as 0.15%. Notably, the same analysts who reject volume estimates still endorse competitive share of voice, which is coherent: a within-set comparison never needed the universe estimate.
"Won't the next model release wipe this out?" Partly, and that is an argument for relative reporting. If a retrieval change drops every brand's absolute rate, the gap between you and a rival is the more stable quantity.
"AI traffic is under 1% of sessions — why are we funding this?" Honest answer: it is a leading indicator of consideration-set membership, not yet a traffic channel. Conductor's US benchmark puts AI referral traffic at 0.48% of total traffic in financial services versus 17.42% from organic Google. The counterweight, per Semrush's 126-million-prompt index, is that 45% of marketing leaders say they cannot accurately measure AI visibility today and only 9% have tools tracking all relevant metrics.
When referral clicks are scarce but presence still matters, read how to quantify and report AI visibility without relying on traffic in Zero-Click AI Answers: Measuring Visibility Without Traffic.
For a practical guide to tying AI visibility and referral signals to measurable sessions and event-level metrics in Google Analytics 4, see how to connect your AI visibility score to GA4.
"Can we just report one score?" Eventually — but compute it last and publish the weights. That is why Geoptimizer's composite is stated openly as EngineScore = 100 × (0.35·MentionRate + 0.25·CitationRate + 0.20·Prominence + 0.20·Sentiment), averaged across ChatGPT, Gemini, Claude and Grok, keeping mention rate and citation rate as separate terms. Competitor tracking runs from 2 rivals on the free plan up to 10 on Pro, so the set you defined in Step 1 can be tracked as a set — the plan limits and scan allowances are listed on the pricing page.
FAQ
What is share of answers, and how is it different from share of voice? Share of voice measures how often a brand appears across a tracked landscape. Share of answers measures how much of the actual AI response a brand owns — whether it is mentioned, recommended, or merely used as an uncredited source. A brand can be present in the citations and absent from the recommendation, which is exactly what happened in the 224 cited-but-not-recommended instances in Lily Ray's AI Overviews study.
How many prompts do I need for a credible competitive analysis? Start around 100 and expect to run 150–200 for a portfolio number with roughly ±5pp precision. Allocation matters more than raw volume: at a fixed budget of 600 answers, 200 prompts × 3 runs gives ±5.4pp while 20 prompts × 30 runs gives ±15.4pp.
Should I run each prompt many times or track more prompts? Both, for different jobs. Use breadth — many prompts, few runs — for the headline share-of-answers number. Use depth — 9–12 runs on 15–25 priority prompts — when you need to diagnose a specific slot, because single runs disagree with a 12-run majority about 19% of the time.
How big does a gap have to be before I report it as real? Attach a margin of error first, discounted by the design effect for repeated runs on the same prompt. As an interpretation rule, category leads under about 1.3 percentage points changed hands routinely in month-over-month tracking, while leads around 2.9 points tended to hold.
Why report per engine instead of one blended score? Because three major models picked the same top brand in only 41.6% of category queries, some engines share only 4–6% of their cited sources with others, and answers carry roughly 3 sources on Gemini versus 15 on ChatGPT. Blending first hides which engine you are actually losing.
Where do competitors' slots usually come from? Overwhelmingly third-party ranked lists. Listicles received 63% of citations in a ~400-million-citation study, 71–86% of them ranked, while brand homepages accounted for just 7% of ChatGPT and 4% of Claude citations. Review platforms then spike at the intent stage, from 1.51% of citations at awareness to 24.27% at intent.
Bringing it together
Five steps: freeze the competitor and prompt sets, choose an allocation and size its margin of error, compute share of answers per engine and per cluster, diagnose the sources supplying your rivals' slots, and set a target against your category's real benchmarks. None of it requires knowing the true size of the prompt universe, which is precisely why it holds up under questioning.
The finding worth carrying into your next planning cycle: 53.7% of tracked categories still have no consistent owner, and traditional SEO strength predicts the winner barely better than a coin flip. That window will not stay open indefinitely.
You can put a baseline on the board this week. Run your domain through Geoptimizer's four-engine AI visibility scanner for a per-engine breakdown across ChatGPT, Gemini, Claude and Grok, with competitors auto-suggested at setup. The free plan tracks 3 AI-suggested prompts and 2 competitors with no card and no expiry — enough to establish whether your category is settled or wide open before you commit a quarter's budget to it.