· 18 min read · Geoptimizer Team

The 2026 B2B SaaS AI Visibility Benchmark, Explained

  • ai-visibility
  • generative-engine-optimization
  • b2b-saas
  • benchmarks
  • measurement
The 2026 B2B SaaS AI Visibility Benchmark, Explained

The most detailed public benchmark of AI visibility in B2B SaaS puts the category average at 56.9 out of 100, with Clio at the top on 89 and LeadSquared at the bottom on 2 — an 87-point spread across 50 companies (DerivateX, The State of AI Visibility in B2B SaaS: 2026 Benchmark Report). But the number most teams should actually care about is the one nobody quotes: the median is 63.5. The mean sits 6.6 points below it, dragged down by a tail of companies that are effectively invisible. So if you scored 57, you are not "average" in any useful sense — you are below the typical company in your category, and the gap is bigger than the headline suggests.

That single distinction changes how you should read every other number in the report. Below is what the benchmark measured, the three ways teams misread their own score against it, and which levers the evidence actually supports — ranked by how strong that evidence is.

What the benchmark actually measured

The study ran 1,400 prompts — seven per company, across four platforms, for all 50 companies — during March and April 2026. The platforms were ChatGPT, Perplexity, Claude and Gemini. Grok is not in the study. Scoring allocates 100 points as mention rate (30), position (30), sentiment (20) and platform breadth (20). The sample spans CRM, SEO tools, project management, payments, field service, property management, security, design and analytics.

Two things deserve saying plainly. First, DerivateX is a GEO agency, so the report doubles as lead generation for the service it sells. Second — and this is the more important half — they published the full 50-company leaderboard with per-component scores, the methodology, and an explicit statement that AI responses are non-deterministic with ±3 to 8 points of run-to-run variance. That is more transparency than most vendor studies offer, and it is why the numbers are worth working with rather than dismissing.

Here is the distribution, which is where the useful signal lives:

Score band Companies Share of sample
81–100 8 16%
61–80 18 36%
41–60 12 24%
21–40 11 22%
0–20 1 2%

Forty-four percent of the sample scored below 50. Fifty-two percent scored above 60. Thirty-two percent cleared 70. And 78% appeared on all four platforms — which means presence is common, but strong presence is not.

The report's own reading of that shape: "The distribution of AI Presence Scores across all 50 companies reveals a bifurcated market, not a normal distribution." Run the arithmetic on the published leaderboard and you get a rough percentile map — 80+ puts you in the top sixth of the sample, 63.5 is the median company, and 56.9 lands somewhere around the 40th–45th percentile. That map is derived from their published table, not a claim the report makes itself, but it is the honest way to convert a score into a position.

Which raises the obvious question: can you compare your own tool's score to any of this?

Three ways teams misread the number

Mostly, no — and the reasons are worth understanding, because each one also tells you something about how to read your internal reporting.

The score isn't portable across tools

DerivateX weights mention 30, position 30, sentiment 20 and platform breadth 20 across ChatGPT, Perplexity, Claude and Gemini. Geoptimizer's AI Visibility Score weights each component differently — mention rate 35, citation rate 25, prominence 20, sentiment 20 — computed per engine and then averaged across ChatGPT, Gemini, Claude and Grok. Two of the four components are different constructs, and one engine is swapped for another. Comparing the raw integers is a category error, not a close-enough approximation.

The direction of the difference is predictable, though, and worth internalising. Any score that gives citation rate real weight will run lower than one that only asks "were you named?" — because being named and being credited are wildly different events. DerivateX's companion study of 233 ChatGPT software recommendations found that only 11.6% cited the vendor's own website; the other 88.4% credited a third party (B2B SaaS AI Citation Study). Their framing is the sharpest sentence in the whole body of research: "The question is not whether AI will cite a source when it recommends you. It will. The question is whether the cited source is yours or a stranger's."

Practically: if your tracker says 42 and the benchmark says the category averages 56.9, your first hypothesis should be methodology, not decline. Check the engine set and the component weights before you check your content calendar.

A fifth of the scale is almost a constant

Forty-four of the 50 companies scored 19 or 20 out of 20 on sentiment. The report is blunt about what that means: "Sentiment is table stakes, not a differentiator. The difference between high and low scorers is not how AI platforms describe them but whether AI platforms mention them at all."

The consequence is arithmetic. If 20 of the 100 points are near-automatic for almost everyone, the effective discriminating range of the scale is closer to 80 points than 100. A 56.9 average on a scale where a fifth of the points are free is a worse result than it sounds, and a five-point improvement earned in the sentiment component is worth far less than five points earned in mention rate. As DerivateX co-founder Apoorv Sharma put it in trade coverage of the study, "Mention rate and position carry 60 of the 100 available points."

There is a related edge case at the bottom of the table. LeadSquared's 2 breaks down as 2/30 mention, 0/30 position, 0/20 sentiment, 0/20 breadth. A company mentioned twice out of thirty opportunities cannot have a meaningfully measured sentiment — that 0/20 is much more plausibly "no observations" than "the models dislike them." Near-zero scores collapse invisible and disliked into the same number, and that is a general property of composite visibility metrics rather than a flaw unique to this one.

Seven prompts is a thin instrument

The benchmark used seven prompts per company. The report concedes ±3 to 8 points of run-to-run variance. Put those two facts together and an eight-point gap between two companies on the leaderboard may not be a real gap at all — it may be the instrument.

This is not a gotcha; it is the same constraint every GEO measurement operates under, and DerivateX's own practitioner guide sets a higher bar than its benchmark does, calling a 20-prompt tracking set "toy-scale" and recommending 50 or more segmented by intent. The lesson transfers directly to your own reporting: your prompt set is the instrument, and it decides your score more than your brand does. We've written a full method for building a buyer-intent prompt set that maps to a funnel rather than to a keyword list, because this is the single easiest place to accidentally measure the wrong thing.

Watch how much work that instrument does in the leaderboard itself.

Brand size does not predict AI visibility

Slack scored 41. Airtable 44, Linear 45, Zoho 46, ClickUp 48 — every one of them below the 56.9 category average. Meanwhile Clio, a legal practice management vendor, topped the list at 89, and AppFolio, in property management, hit 71. Clio outscored Notion (81), monday.com (79) and Salesforce (77).

If your mental model is "big brand, big AI visibility," that row of numbers should end it. Slack's component breakdown is instructive: 9/30 mention, 9/30 position, a perfect 20/20 sentiment, and just 3/20 platform breadth — it was absent from Perplexity entirely in this study. Nobody thinks Slack has an awareness problem. It has a retrieval problem on prompts where it competes against Teams, Discord and Zoom in a crowded horizontal category.

Clio, by contrast, plays a category small enough that it defines it. As the report puts it: "The top scorers (Clio, Procore, Loom, Figma) are not simply participants in their categories. They are the names AI platforms use to define the category itself."

The other structural finding is that spread within categories beats spread between them. Direct competitor pairs diverged by 4 to 27 points:

Category Pair Gap
Field service ServiceTitan 68 vs Jobber 41 27
Payments Stripe 65 vs Razorpay 39 26
Automation Zapier 63 vs Make 40 23
Product analytics Mixpanel 68 vs Amplitude 49 19
SEO analytics Ahrefs 83 vs Semrush 68 15
Enterprise SEO Conductor 53 vs BrightEdge 39 14
Property management AppFolio 71 vs Buildium 67 4

Twenty-seven points separating two field-service platforms is not a category effect — both companies are answering the same buyer prompts on the same four engines. It is a citation-surface effect, and it is addressable. Which is why the useful benchmark is never the category mean; it is the competitor one row above you. That is a different measurement job, and it needs a share-of-answers method rather than a share-of-voice one — you want to know which specific prompts your rival owns and which sources are winning those slots for them.

Read your components, not your total

A single composite number tells you where you stand. The component breakdown tells you what to do on Monday. Here is a diagnostic that maps the pattern to the problem:

  • High sentiment, low mention (say, 19–20/20 sentiment with mention at or below 8/30) — you have a distribution problem, not a reputation problem. The models have nothing bad to say about you; you simply do not come up. The benchmark identified ten companies in exactly this state, including Close (1/30 mention), Freshworks (5/30), BrightEdge (6/30) and Toast (8/30). The fix is presence in the sources engines pull from, not messaging work.
  • Decent mention, poor platform breadth — you are winning one or two engines and missing the others. Claude is the hard one: it mentioned 88% of tested brands versus 100% for both ChatGPT and Gemini. Companies absent from Claude in the study included Zapier, Mindbody, Toast, Razorpay, WebEngage and Chargebee.
  • Named often, never cited to your own domain — a citation-surface problem. This is the default state rather than the exception, given that 88.4% of citations go to third parties.
  • Mentioned but consistently ranked third or fourth — a position problem. Position is worth 30 points in this benchmark for a reason, and the top of the table wins on consistency: "All 10 of the top 10 scorers hold position #1 on all 4 platforms."

The contrast between the ends of the distribution makes the point numerically. High scorers (26 companies at 60 or above) averaged 18.8/30 on mention rate and 18.4/20 on platform breadth, with 92% present on all four platforms. Low scorers (8 companies at 35 or below) averaged 3.0/30 on mention and 2.5/20 on breadth. The gap is not sentiment, and it is not position among people who already mention you. It is whether you show up at all.

Which levers move fastest, ranked by evidence strength

Now the part every marketing lead actually wants: given a component-level diagnosis, what moves the number? Ranked honestly by how much evidence stands behind each.

Strong: your off-domain mention footprint. Ahrefs studied 75,000 brands and found branded web mentions the single strongest correlate of AI Overview visibility at Spearman ρ = 0.664 — well ahead of Domain Rating (0.326), referring domains (0.295) and backlinks (0.218) (Ahrefs brand correlation study). Their conclusion: "Brand's presence across the web—not just your own website—is what AI Overviews draw on when deciding whether to mention you." Ahrefs is explicit that correlation is not causation and that all these coefficients are moderate-to-weak, so treat it as the best available signal rather than proof. DerivateX's guide puts a rough number on the same idea from the practitioner side: roughly 80% of what determines AI citation lives off your domain.

Strong: being inside third-party comparison content. This follows directly from the 88.4% figure. If most of your citations will be somebody else's page, the highest-leverage work is getting into other people's roundups rather than publishing your own /vs pages and hoping. We've broken down the pitching mechanics separately in how to get into the "best tools" lists AI engines cite.

Moderate: the structural format of the pages that do cite you. DerivateX audited 143 citation-winning pages: 100% used numbered or bulleted lists, 78% carried the current year in the title, 68% included comparison tables, 56% had FAQ sections, and 57% had all three of list, table and FAQ. That is a pattern in what gets cited, not a controlled test — but it is cheap to act on, and format changes are a rewrite rather than a new content programme. In one practitioner survey of 169 marketers, 14 respondents credited format changes rather than new content as their primary AI visibility win (CommonMind).

Contested: review platforms. The evidence genuinely conflicts here, and you should know that before you fund anything. DerivateX found review aggregators (G2, Capterra, TrustRadius) at just 0.9% of citations in ChatGPT software recommendations. But Peec AI's analysis of 30 million cited sources across five engines puts G2 at #6 among all cited domains and inside Perplexity's top five. And Seer Interactive, studying 804,491 AI responses across 1,926 brands, found brands with no Trustpilot profile had a 1% median citation rate while brands with even a minimal profile of 1–13 reviews hit 53.5%. The best reconciliation the data supports is that this is engine- and vertical-dependent — ChatGPT leans on editorial content, Perplexity leans harder on review and community domains — and Seer's result concerns Trustpilot in largely consumer verticals, not G2 in B2B SaaS. Measure it in your own category rather than treating any of the three as settled.

Weak: llms.txt. Ahrefs checked 137,210 domains in May 2026 and found 28% had published a valid llms.txt — then found that 97% of those files got no traffic at all that month. Of the requests that did land, only 19.5% came from AI-related bots (Ahrefs llms.txt study). Ship it if it takes twenty minutes; do not put it on a roadmap slide as a growth lever. Crawler access in robots.txt is a genuinely different matter and does affect whether engines can read you — keep the two separate in your planning.

Weak: publishing more blog posts. DerivateX names "more blog content is the answer" as practitioner mistake #1, sizing on-site work at roughly 20% of the lever. That does not mean stop publishing; it means stop expecting volume alone to move a mention rate.

One more reason not to reach for your usual SEO playbook: AI citation and Google ranking are largely decoupled. Across 15,000 long-tail queries, only 12% of AI-cited URLs ranked in Google's top 10 for the same prompt — Perplexity was highest at 28.6%, ChatGPT's in-text citations lowest at 8%. Ranking well is not a reliable proxy for being cited.

How long it takes, and how to tell it worked

Whatever you change, the measurement problem arrives immediately: the underlying signal is noisy enough that a single scan can't tell you anything.

MaxAEO ran 1,247 buyer-intent prompts across 52 B2B software categories daily on eight platforms from 1 March to 29 May 2026 — 897,840 answers, plus a control that re-ran 200 prompts ten times within the same hour (AI answer volatility study). The findings that matter for your reporting cadence:

  • 17% of prompts return a different set of recommended brands than the day before, and the median brand list holds for just 5 days.
  • Roughly one-third of that day-over-day churn is sampling noise, per the control test — the same prompt, the same hour, a different answer.
  • Stability varies enormously by engine: Claude churns at 9% and holds a brand list for a median 11 days, while Perplexity churns at 27% with a median of 3 days. Grok sits at 21%.
  • An average of 2.3 brands per prompt appeared in at least 80% of daily snapshots — the anchors — with four to nine challengers rotating through the remaining slots.

That last number is the strategic one. Most categories have two or three brands the engines treat as fixtures and a rotating cast fighting for the rest. Moving from challenger to anchor is the actual goal, and you cannot see it happening in a single scan.

The practical consequence is a measurement discipline, not a tactic: freeze your prompt set, report per engine, and read a rolling window with a confidence band instead of a point estimate. Geoptimizer's headline score is a 7-day rolling window with a band around it for exactly this reason, and a single on-demand scan is labelled a snapshot rather than a score — which is the honest description of what one run of a non-deterministic system produces. Report every engine you track, too — dropping one changes your number as surely as dropping a prompt does.

On timelines: the practitioner playbooks describe 90-day turnarounds in weak categories with a clear ICP niche, and 12 to 18 months for horizontal categories with entrenched incumbents. Those are agency case studies without independent verification, so treat them as shape rather than promise. The shape is right, though — mention rate moves slower than anything you are used to from paid search.

What being below the line actually costs

It would be easy to file all of this under "interesting, not urgent," especially since AI referral traffic is still a rounding error for most sites. AI Mode was used in just 0.34% of US searches between January and April 2026 (SparkToro/Similarweb clickstream). That is a fair objection, and the honest answer is that the channel is small but the influence is not — and it happens inside the answer, not in the click.

The buyer data is where this stops being abstract. G2's 2026 buyer research found that 51% of B2B software buyers now start research with an AI chatbot more often than with Google — up from 29% a year earlier. 69% chose a different vendor than they had planned based on chatbot guidance. And one-third bought from a vendor they had never heard of before. In G2's fuller report, 82% of buyers had sourced software recommendations from an AI chatbot in the previous 24 months, and evaluation is now the longest stage of the journey for 40% of them.

Read those three numbers together and the stakes are clear. A third of buyers are willing to purchase from a vendor they had never heard of, which means an unknown competitor can enter your shortlist on the strength of a citation. Sixty-nine percent switched intent based on what the assistant said, which means being absent from the answer costs you deals you would previously have won on brand alone. For a company scoring in that bottom 44%, the loss is not a traffic line item — it is shortlist slots, disappearing before anyone fills in a form.

One final caution against over-reading any single benchmark. Three credible 2026 studies report three different B2B SaaS averages: 56.9 (DerivateX, 50 companies), a median of 62 (Foglift, 1,148 SaaS brands, Q1 2026), and 50 (Mojo Dojo, 712 B2B companies, with 89% of them scoring below 70). Different engine sets, different components, different samples. A benchmark tells you the shape of the distribution — bifurcated, with mention rate doing most of the work and a long invisible tail — not your grade.

So use 56.9 as a reference point and never as a target. Your real target is the competitor one row above you in your own category, measured on your own frozen prompt set, on every engine your buyers actually use. If you want that number for your brand today, run a scan across ChatGPT, Gemini, Claude and Grok — all four engines are included on every Geoptimizer plan, including the free one — and see which component, mention, citation, prominence or sentiment, is holding your score down. That component, not the category average, is your roadmap.

FAQ

What is a good AI visibility score for a B2B SaaS company in 2026? Against the DerivateX benchmark, the median company scored 63.5 and the top 16% scored above 80, so anything north of 70 is genuinely strong and anything below 50 puts you in the bottom 44%. Treat those bands as a rough map rather than a grade, because other 2026 benchmarks report B2B SaaS averages of 50 and a median of 62 using different components and engine sets.

Why is the category average 56.9 but the median 63.5? Because the distribution is bimodal rather than normal. A cluster of near-invisible companies — 11 scored between 21 and 40, and one scored 2 — pulls the mean down 6.6 points below the median. That is why "we're around average" is a misleading self-assessment: the average is below the typical company.

Can I compare my GEO tool's score to the 56.9 benchmark? Not directly. The benchmark weights mention, position, sentiment and platform breadth across ChatGPT, Perplexity, Claude and Gemini. Most trackers use different components — a citation-rate component in particular will push a score lower, since only 11.6% of ChatGPT software recommendations cite the vendor's own site — and a different engine mix. Compare against your own history and your named competitors instead.

Does a bigger brand mean better AI visibility? No. In the benchmark, Slack scored 41, Airtable 44 and ClickUp 48 — all below the category average — while Clio hit 89 and AppFolio 71. Vertical category-definers consistently outperformed much larger horizontal brands, because the prompts they compete on have fewer credible answers.

How often should I measure AI visibility? Weekly, as a rolling window with a confidence band — never as a single scan. About 17% of prompts return a different brand set day-over-day, roughly a third of that is pure sampling noise, and per-engine stability ranges from 9% daily churn on Claude to 27% on Perplexity. A one-off scan is a snapshot; a trend across a frozen prompt set is a measurement.

Is llms.txt worth implementing for AI visibility? It is cheap, so ship it if you like, but do not budget for it as a lever. Of 137,210 domains checked in May 2026, 28% had published a valid llms.txt and 97% of those files got no traffic at all that month. Crawler access in robots.txt is the part of technical GEO that genuinely affects whether engines can read you.

Keep reading

See it on your own domain.

Free visibility check across ChatGPT, Gemini, Claude, and Grok — about 30 seconds.

Run the free check