· 22 min read · Geoptimizer Team
How to Get Into the Best Tools Lists AI Engines Cite
- generative-engine-optimization
- ai-citations
- listicles
- b2b-saas
- off-site-geo
When ChatGPT recommends a B2B software tool, it cites that tool's own website only 11.6% of the time — the other 88.4% of the credit goes to somebody else's page, according to a study of 233 recommendations across 40 B2B SaaS categories. Getting into the "best X software" lists AI engines cite is therefore mostly off-site work: find the specific URLs each engine already retrieves for your category prompts, qualify which of those lists are actively maintained, and pitch the editor on the gap your product fills. The rest of this guide is how to do that engine by engine, and how to tell whether it worked without fooling yourself with one noisy scan.
If your category prompts return three competitors and no mention of you, the reflex is to go rewrite your homepage. That reflex is aimed at the wrong document.
The uncomfortable arithmetic behind a "best X software" answer
Start with the format data, because it explains almost everything else. Across roughly 25,000 unique most-cited URLs pulled from six LLMs — ChatGPT, Copilot, Gemini, Google AI Mode, AI Overviews and Perplexity — in March and April 2026, about half were listicles, and 63% of nearly 400 million citations pointed to listicles. Between 71% and 86% of those listicles were ranked Top-N lists. Listicles made up 40–65% of most-cited URLs depending on the model, Copilot at the low end and Gemini at the high end.
That is all query types. Narrow to commercial intent and the effect sharpens: Wix Studio's AI Search Lab analysed 75,000 AI answers and over a million citations across ChatGPT, Google AI Mode and Perplexity, and found listicles took 40.9% of citations on commercial-intent queries versus 21.7% on informational ones. Their summary line is worth pinning above your desk: "Query intent — not industry or model — most strongly predicts which content gets cited."
Narrow again to B2B software and the pattern holds from three independent directions:
- Across 26,283 source URLs from 750 top-of-funnel prompts on ChatGPT, 43.8% of all cited page types were "best X" blog lists.
- Across 1,260 solution-aware prompts spanning 250 niche B2B software categories, run over one business week in June 2026 on four engines, 70.8% of citations pointed to URLs using "best" or similar language and 51.6% carried "2026" in the title.
- In a study of 21,075 AI responses and 25,337 citations across eight brands and five engines from 14 April to 13 July 2026, one B2B technology services brand saw listicles account for 61% of its citation share while its own domain held 0% of its top-30 cited URLs.
Two things follow. The answer to "best X software" is assembled from other people's ranked lists, not from your product page. The pool of pages an engine will even consider is also small — and it is not the Google SERP. Ahrefs ran 15,000 long-tail queries through ChatGPT, Gemini, Copilot and Perplexity and found only 12% of AI-cited URLs ranked in Google's top 10 for the prompt that produced them, with roughly 80% not ranking anywhere in Google for that query. You cannot infer your citation set from your rank tracker.
Is it worth the effort? A March 2026 survey of 1,076 B2B buyers found 51% now start software research with an AI chatbot more often than with Google, up from 29% in April 2025; 69% picked a different vendor than planned because of chatbot guidance, and 33% bought from a vendor they had never heard of. That research is G2's, and G2 sells presence on a review platform. Most studies cited here come from companies selling adjacent tools or services, ours included — worth saying once so you can weigh the denominators yourself.
What "invisible" actually looks like
A benchmark of 50 B2B SaaS companies against 1,400 buyer-intent prompts on ChatGPT, Perplexity, Claude and Gemini in early 2026 produced an average AI presence score of 56.9 out of 100 and 44% of companies scoring below 50 — "functionally invisible". Scores ranged from 2 to 89, with wide gaps inside the same category: ServiceTitan 68 versus Jobber 41, Zapier 63 versus Make 40.
The diagnostic detail matters more than the ranking. Ten of the fifty had perfect sentiment alongside mention rates of 8 out of 30 or lower. The models had nothing bad to say about them; they simply never came up. That is a frequency problem on other people's pages, not a perception problem on yours — and messaging work on your own site will not fix it.
The four doors, and why each engine opens a different one
There is no single "AI listicle strategy," because the engines do not share a citation pool. Across 3.7 million cited URLs collected from Q3 2025 through Q1 2026, 91.07% appeared in only one engine and just 2.37% appeared in all three engines measured for the same prompt. Scope: ChatGPT, Perplexity and Google AI Overviews only, 20,000 prompts queried twice daily plus three 5,000-prompt cohorts, European markets with a Spain-heavy sample. The single-engine share eased to roughly 88% by Q1 2026 — still high enough that a placement earning ChatGPT citations may do nothing on Gemini.
Four broad doors lead into the pool, and their value swings hard by engine.
| Surface | Where it carries weight | Evidence |
|---|---|---|
| Third-party editorial listicles | Every engine, strongest on commercial-intent prompts | 40.9% of commercial-intent citations; 80.9% of listicle citations in professional services go to third-party rather than self-promotional lists |
| Review networks (G2, Capterra, GetApp, Software Advice) | Perplexity and Google surfaces more than ChatGPT | G2 network = 8.0% of B2B software citations; a ChatGPT-only study logged zero G2 and Capterra citations |
| Communities (Reddit, LinkedIn) | ChatGPT and Grok for Reddit; all engines for LinkedIn on professional queries | Reddit is 16.7% of ChatGPT citations and 16.3% of Grok's — and 0% of Claude's on SaaS/tech queries |
| Video (YouTube) | Google surfaces and Grok | 7.4% of AI Overview citations in B2B; 15.1% of Grok's citations |
The Reddit contrast is the cleanest illustration. Ahrefs' Brand Radar data puts reddit.com at 16.7% of ChatGPT citations across US queries in July 2026, ahead of Wikipedia at 8.9% and Forbes at 3.3%. On Grok, across 1.9 million US queries in June 2026, reddit.com takes 16.3%, youtube.com 15.1% and facebook.com 13.9%, with no B2B publication in the top ten. Otterly's June 2026 analysis of 379,321 Claude citations on SaaS and technology queries recorded brand domains at 64.0%, community and forum at 1.3%, and Reddit at exactly zero. Same category of question, three completely different source diets.
Review networks deserve their own note, because the data conflicts. In the 250-category B2B study, the G2 network accounted for 8.0% of citations with G2.com alone at 5.8% — the single most-cited domain — and roughly 71% of G2's citations came from Perplexity. The 40-category ChatGPT-only study found review aggregators at 0.9% of citations, with G2 and Capterra each receiving zero. Both can be true: different engines, different prompt sets, different denominators.
What review networks reliably are is an inclusion gate, not a ranking lever. Across high-intent "X alternatives" prompts, 100% of the tools ChatGPT named had Capterra reviews and 99% had G2 reviews — but the correlation between review counts and position was weak and slightly negative (Capterra reviews −0.21, G2 reviews −0.16, G2 score −0.11). Presence is table stakes; grinding review counts from 200 to 400 is not what moves you up the answer.
One structural change to plan around: G2's acquisition of Capterra, Software Advice and GetApp closed in February 2026. Modelling across 25,755 citations from 200 buyer-intent prompts on five engines suggests the combined network's share of bottom-of-funnel B2B citations rises from 2.09% to 3.68%, moving it from fourth to second most-cited source. Fewer, larger gates.
LinkedIn is the surface B2B teams most often under-weight. Analysis of 1.4 million citations found LinkedIn is the single most-cited domain for professional queries across six engines, climbing on ChatGPT from roughly rank #11 in November 2025 to #5 by February 2026. The internal mix shifted too: profiles fell from 33.9% to 14.5% of LinkedIn citations while posts rose from 20.9% to 26.0% and long-form articles from 6.0% to 8.9%. Your team's posts are now more citable than your company page.
Step 1 — Build a citation set, not a mention report
Most GEO reporting answers "was my brand mentioned?" For this work that is the wrong first question. The one that produces a plan is: which URLs did each engine cite when it answered my category prompt?
Run 20–40 buyer-intent prompts for your category on each engine with web search enabled, and log the cited URLs — not just the brands named. Lists that recur across prompts and engines are your priority targets, because they are already inside the retrieval pipeline. A list that never surfaces in any answer is a link-building target at best.
Do this per engine and do not average. Beyond the 91% single-engine finding, the composition of top-cited domains differs in ways that change your tactics: in the June 2026 B2B study, roughly 75% of the top 100 domains on AI Overviews and Gemini were vendor sites, versus roughly 25% on ChatGPT and Perplexity. If Gemini is your priority engine, your own pages carry more of the load; if ChatGPT is, third-party editorial carries it.
Three practical notes on running the audit:
- Answer depth and concentration vary. Semrush's index of 126 million US AI search prompts from January to April 2026 found ChatGPT averages 15 sources per response while Gemini averages 3. And the pool is wide but shallow — across ~730,000 US ChatGPT conversations, the top 10 domains accounted for only 12% of all citations. You are not chasing ten publishers; you are chasing the ten that recur in your category.
- Audit in fresh chats. That same conversation study found citations triggered on 12.6% of turn-one messages versus 4.5% at turn 10 and 3.0% at turn 20, with about six unique citations per conversation. Deep in a thread, engines stop sourcing.
- Separate mentions from citations. On Gemini, the overlap between brands mentioned and domains cited is only 30%. Being named and being linked are different outcomes and need different numbers.
That last split is what Geoptimizer's AI Visibility Score is built around: mention rate carries 35% of the weight and citation rate a separate 25%, computed per engine and then averaged, with the formula published rather than described. For this job the per-engine citation breakdown is the output you want, and the free Prompt Ideas Generator will turn your domain into 20 buyer-intent prompts to seed the audit if you don't already have a prompt set you trust. How to build a set worth freezing — funnel allocation, branded versus unbranded, wording sources — is covered in How to Choose the 25 Prompts You Track, and finding the specific third-party URLs that keep supplying your rivals' answer slots is the source-gap step of Competitive AI Visibility Analysis.
Step 2 — Qualify the lists worth pitching
A recurring cited URL is a candidate, not a target. Four filters, in order.
Does it feature your direct competitors? A list that already ranks three rivals is one where your absence is conspicuous and your inclusion is editorially defensible.
Is it maintained? Of 1,100 analysed cited blog lists, 79.1% had been updated in 2025, 26% within the previous two months, and 57.1% had been modified after first publication. Check the Wayback Machine for real edits, not a changed date. But don't over-rotate on recency: in Ahrefs' analysis of ChatGPT retrieval, the median cited search-result page was around 500 days old, and retrieved-but-uncited pages skewed younger. Inside the retrieval set, relevance beats recency. "Maintained" is the filter; "published last month" is not.
Does the title match how buyers phrase the query? The same analysis found title-to-query semantic match separates cited pages from merely retrieved ones: cited pages averaged 0.602 cosine similarity to ChatGPT's fan-out queries versus 0.484 for retrieved-but-uncited pages, and readable URL slugs were cited 89.78% of the time versus 81.11% for opaque ones. "Best [category] software for [segment] (2026)" is a better host than "12 tools we like."
Is it a real publication? In the June 2026 B2B study, 8.6% of citations went to spam or synthetic sites — 14.5% of ChatGPT's and 11.3% of Perplexity's, versus 0.16% for AI Overviews and 0.46% for Gemini. Independent academic auditing of 712 real queries across four engines found roughly 16% of cited sources showed evidence of being AI-generated. A placement on a synthetic listicle farm is visibility rented on ground the engines are working to clean up.
Aim for 20–40 qualified targets. That is a quarter's worth of outreach, not a week's.
Step 3 — The outreach that actually lands
Pitch the gap, not the product
The editorial principle in Position Digital's listicle outreach guide is blunt and correct: editors do not care that you exist, they care whether including you improves the list. Lead with what the list is missing — a segment it doesn't cover, a use case where its current picks fit poorly, a pricing tier absent from the lineup.
Their process: identify the gap, find the in-house editor rather than a freelance writer, personalise to specific elements of the article, hand over a paste-ready 40–60 word blurb in the publication's voice, then follow up two to three times with new information each time. They report 18 placements in seven months, including one that landed after 14 follow-ups over five months. A separate practitioner benchmark puts the placement rate for well-targeted pitches at 15–25% — so 30 qualified pitches at 20% is six placements a quarter. That is a plan. "Get on the lists" is not.
Control the sentence next to your name
The paste-ready blurb is not a courtesy, it's the payload. Engines extract descriptions from the list itself, and Wix Studio found 44.2% of LLM citations are extracted from the first 30% of a document, with 31.1% from the middle third and 24.7% from the final third. Write the differentiator and the "best for X" qualifier yourself, in the publication's register, so the sentence that gets lifted is the one you'd have chosen.
Negotiate position, not just presence
Peec AI's study of nearly 200,000 AI responses across eight engines between September 2025 and March 2026 found that in B2B SaaS, brands at rank #1 in frequently-cited third-party listicles showed +16.5 percentage points of visibility versus baseline and appeared roughly 1.17 positions earlier in AI answers.
Treat that number carefully. Peec states plainly that this is "observational research, not a randomized experiment," that the results describe "associations," and that they apply to "frequently retrieved listicles, not every listicle on the web." Brands sitting at #1 also tend to have more of everything else, and nobody has shown that moving from #6 to #1 causes a 16.5-point lift. What the data supports is a priority ordering: argue for a higher slot and a category-specific "best for" label rather than a generic mention at the bottom. The same study reports diminishing returns after a few repeated placements, so breadth across distinct lists beats stacking one publisher.
One adjacent asymmetry worth exploiting: in the DeltaV study, comparison pages had the highest citation rate per retrieval at 1.87 despite representing only 4.1% of citation share. Rarely retrieved, heavily cited when they are — and third-party "X vs Y" pages and your own honest comparison pages both qualify.
Work community and video surfaces deliberately
Reddit is the surface most likely to be misread. ChatGPT retrieves it constantly and credits it rarely: Reddit pages were cited only 1.93% of the time when retrieved, and accounted for 67.8% of all retrieved-but-uncited URLs, while ordinary search-result URLs were cited 88.46% of the time. As Search Engine Journal summarised that dataset, ChatGPT "is using Reddit extensively to understand topics, gauge consensus, and build context — but it almost never gives Reddit the credit." Reddit influence therefore shows up in your mention rate more than your citation rate. Its role is narrowing, too: overall Reddit citation share halved from 2.02% to 1.01% between October 2025 and January 2026, while the share of responses where Reddit was the only cited source rose 31%.
For video, YouTube mentions showed the highest correlation with AI visibility across 75,000 brands with Domain Rating above 40 — a Spearman coefficient of about 0.737, ahead of branded web mentions at 0.664 on ChatGPT, with raw backlink counts weak. Ahrefs adds the necessary caveat that correlation isn't causation. The consistent pattern across these datasets: third-party mentions track AI visibility far more closely than links do.
Where this goes wrong
Buying your way in. Google's spam policies define link spam as creating links "primarily for the purpose of manipulating search rankings," explicitly including "exchanging money for links, or posts that contain links" and "exchanging goods or services for links"; paid arrangements require rel="sponsored" or rel="nofollow". The site reputation abuse policy covers third-party content published on a host site mainly to exploit that host's ranking signals. Reciprocal list swaps and paid placement marketplaces sit close to both lines.
Publishing your own "best" list carelessly. The evidence points both ways. Ahrefs' data supports publishing them — for software recommendations, first-party cited pages broke down as 37.2% landing pages, 34% blog lists and 15.6% homepages. But third-party lists took 80.9% of listicle citations in professional services, and Lily Ray's early-2026 analysis found multiple SaaS brands lost 30–50% of organic visibility after the December 2025 Google core update, concentrated in blog sections full of self-promotional "best" listicles — many of which had only been "updated" by adding 2026 to the title. Google has not confirmed a targeted update, and Ray notes the drops "will also impact visibility across other LLMs that leverage Google's search results."
The defensible middle is the one Ahrefs applies to itself: keep such lists under 0.5% of blog content, don't always rank yourself first, make attribution explicit, link to alternatives. As Wil Reynolds put it in that write-up, "Any marketer worth a damn knows listing yourself at top doesn't build trust with buyers."
Expecting ads to substitute for citations. OpenAI launched ChatGPT ads on 9 February 2026 with wider rollout in May. In one tracked sample, ads appeared in 77% of responses but the advertiser was present in the organic answer only 5% of the time, with B2B software the largest advertiser vertical at 21.7% of impressions. Ads sit beside the recommendation; they don't become it.
How to know it worked
Give it weeks, not days. Practitioners report first citations 2–6 weeks after placement, with 90 days before judging ROI. A longitudinal study of 81 newly published pages found 34 cited by ChatGPT within 30 days while 47 (58%) were never cited in the window, with citation odds falling sharply after 60 days uncited.
Don't try to prove it with one scan. A July 2026 variance-components study of 12,933 responses across GPT-5.2, Gemini 3 Flash and Perplexity found within-prompt resampling accounted for 34.8% of variance and query language 26.5%, while brand identity — the thing you are measuring — accounted for 1.5% of single-response variance. Its recommendation is counterintuitive and useful: don't repeat the same prompt more than about five times, and spend the query budget on paraphrases, languages and models instead.
That is the same statistical reality behind why two GEO trackers report different scores for the same brand, and it's why a defensible setup reports a rolling window with a confidence band rather than a number from one run. In Geoptimizer that means a 7-day rolling window with a stated band and on-demand scans labelled as snapshots — a spot-check at two weeks tells you whether anything moved, a full sweep across all four engines at 90 days tells you whether it held. Because every plan includes ChatGPT, Gemini, Claude and Grok rather than selling engines as add-ons, you can see which engine a placement actually moved — given how little the citation pools overlap, that is the only comparison that means anything.
Track mentions, citations and links as three metrics. A placement can lift your mention rate without touching citation rate, or vice versa. Links are now a third outcome: on 7 May 2026 ChatGPT began embedding inline brand links, and the share of answers containing a brand link jumped from about 0.43% to 6.20% overnight — roughly 2% to 29% among brand-mentioning answers, with about 79% pointing to homepages. No comparable shift on Perplexity, Gemini or Copilot.
Expect an upper bound. Semrush's clickstream analysis of over a billion lines of US data found web search was triggered on 34.5% of ChatGPT queries in February 2026, down from 46% in late 2024. Roughly two-thirds of answers involve no live retrieval, so off-site placements move the grounded subset directly and the parametric subset only slowly. It's also why a score can shift for reasons unrelated to your work — separating a real visibility change from a new model release is a prerequisite for reading any placement result honestly.
A 30/60/90 plan
Days 1–30 — audit and target list. Assemble 20–40 buyer-intent prompts. Run them on each engine in fresh sessions with web search enabled, logging cited URLs per engine rather than just whether you were named. Baseline mention rate and citation rate separately, per engine, over at least a week. Output: a ranked list of 20–40 target URLs, each tagged with the engines that cite it and whether it features your competitors.
Days 31–60 — pitch and clear the gates. Send 25–30 personalised pitches to in-house editors, each leading with that list's specific gap and carrying a paste-ready blurb. In parallel, complete your G2-network profiles — presence there is near-universal among tools ChatGPT names. Publish or refresh two or three honest comparison pages. If LinkedIn matters in your category, shift effort from the company page to posts and long-form articles by named people.
Days 61–90 — re-scan and reallocate. Re-run the same frozen prompt set on all four engines and compare per engine, never in aggregate. Ask three questions: did the new lists appear as cited URLs, did your mention rate move on the engines citing those lists, and did anything move on engines that don't. Double down on publications that produced citations; drop the ones that produced a backlink and nothing else. Then repeat — the cited-domain set is not stable, as LinkedIn's climb from #11 to #5 and Reddit's halving both show.
One honesty note: off-site work is the highest-leverage route on average, not the only route. A published eight-month engagement with video platform Gumlet took the company from invisible for "best video hosting platforms" to roughly 20% of inbound revenue from AI-discovery users at 2.3× the conversion rate of traditional search — via owned-content work: 17 pages, entity mapping, money-page repositioning. If your engine mix leans toward Gemini and AI Overviews, where roughly 75% of top-cited domains are vendor sites, your own pages carry more of the load than this post's thesis implies. Audit first; the answer is category-specific.
FAQ
How do I find which listicles ChatGPT cites for my category? Run 20–40 buyer-intent prompts in fresh chats with web search enabled, and record every cited URL rather than just noting whether your brand appeared. Lists recurring across several prompts are already inside the retrieval pipeline and are your priority targets. Repeat on Gemini, Claude and Grok separately — 91.07% of cited URLs in one large study appeared in only one engine, so a ChatGPT-derived target list will not transfer.
Do G2 and Capterra reviews get me into AI answers? They act as an inclusion gate rather than a ranking lever. Across high-intent "alternatives" prompts, 100% of tools ChatGPT named had Capterra reviews and 99% had G2 reviews, yet the correlation between review counts and position was weak and slightly negative. Citation weight also varies by engine: the G2 network took 8.0% of citations in one 250-category B2B study, around 71% of that from Perplexity, while a ChatGPT-only study recorded zero. HubSpot's operational guidance is 50+ reviews at a 4.0+ average — treat that as a floor, not a strategy.
How long does a new listicle placement take to show up in AI answers? Typically 2–6 weeks to first citation, with 90 days as the point to judge return. A longitudinal study of 81 new pages found 34 cited by ChatGPT within 30 days and 58% never cited within the tracking window, with odds dropping sharply after 60 days uncited.
Should we publish our own "best tools" list? Cautiously. First-party lists do get cited — 34% of first-party cited pages for software recommendations were blog lists — but third-party lists took 80.9% of listicle citations in professional services, and several SaaS brands lost 30–50% of organic visibility after the December 2025 core update in blog sections dominated by self-promotional lists. The defensible pattern is Ahrefs' own: under 0.5% of blog content, don't always rank yourself first, disclose the affiliation, link to genuine alternatives.
Why did my score barely move after a placement? Three measurable reasons. Roughly two-thirds of ChatGPT queries trigger no live web search, so placements move only the grounded share of answers. Response-level noise is large — resampling alone drove 34.8% of variance versus 1.5% for brand identity — so a single before/after comparison cannot detect a real effect. And the placement may have moved your mention rate without moving citation rate, which is invisible if you track only one.
Start with the citation set
The shift this asks for is small but load-bearing: stop asking "am I mentioned?" and start asking "which URLs got cited, on which engine, for which prompt?" That question produces a target list. The target list produces outreach. The outreach produces placements you can measure — slowly, per engine, with a confidence band rather than a single number.
To get the baseline without setting anything up, run the free AI Visibility Check to see whether ChatGPT, Gemini, Claude and Grok currently mention your brand, then use the Prompt Ideas Generator to build the 20 buyer-intent prompts your citation audit should start from. The Free plan tracks three AI-suggested prompts across all four engines indefinitely — enough to learn which door your category opens through before you spend a quarter on outreach.