· 22 min read · Geoptimizer Team

Do Backlinks Matter for AI Citations? The 2026 Data

  • generative-engine-optimization
  • ai-citations
  • backlinks
  • digital-pr
  • measurement

Yes — but as a threshold you clear once, not a dial you keep turning. In Ahrefs' study of 75,000 brands, branded web mentions correlated with AI Overview visibility at 0.664 (Spearman), while the raw number of backlinks came in at 0.218. Yet in the same period, SE Ranking analysed 129,000 domains and found referring domains to be the single strongest predictor of ChatGPT citations. Both results are real, and once you see why they don't conflict, the budget question answers itself: links buy you a place in the candidate set, mentions decide whether you're the answer.

That distinction matters right now because most teams are stuck between the two findings. The State of Link Building 2026 survey of 500 SEO professionals found that 74% believe backlinks affect AI visibility, only 24% are tracking AI visibility at all, and just 19% have changed how they build links because of it. A strong belief, almost no measurement, and virtually no behaviour change. If you're spending $3,000 a month or more on links — as 64% of that sample are — you're funding a hypothesis you haven't tested. So let's test it.

The two studies that look like a contradiction

Start with the number everybody quotes. Ahrefs' AI Overview brand correlation study took 75,000 brands — filtered to domains with Domain Rating above 40, on keywords with at least 800 monthly searches — and ranked familiar metrics by their Spearman correlation with AI Overview brand visibility:

Factor Spearman correlation
Branded web mentions 0.664
Branded anchors 0.527
Branded search volume 0.392
Domain Rating 0.326
Referring domains 0.295
Branded traffic 0.274
Number of backlinks 0.218

Ahrefs are careful about what this shows. Their write-up states plainly that "correlation ≠ causation," and notes that "all the factors we studied revealed moderate to very weak correlations on the Spearman scale" — an unusually honest framing for a study whose headline now circulates as proof that links are finished.

Their multi-platform follow-up across ChatGPT, AI Mode and AI Overviews extended the same sample and turned up something stranger: YouTube mentions — a brand name in a video title, transcript or description — correlated at roughly 0.737, the strongest single factor measured on any platform. The engines also pulled in different directions. AI Mode leaned hardest on traditional brand-authority signals (branded search volume at 0.466), while ChatGPT correlated most weakly with established authority metrics — which Ahrefs read as ChatGPT being a more accessible entry point for emerging brands.

That spread shows up again in an earlier Ahrefs analysis of the top 50 Brand Radar domains, built on June 2025 data covering roughly 76.7 million AI Overviews, 957,000 ChatGPT prompts and 953,500 Perplexity prompts. The mentions-to-visibility correlation was 0.65 on Google AI Overviews, 0.30 on Perplexity, and 0.15 on ChatGPT. So "mentions beat links" is itself an AI-Overviews-flavoured claim; on ChatGPT, in that sample, mentions barely moved the needle.

Now the other side. SE Ranking's ChatGPT citation study ran XGBoost regression with SHAP analysis over 129,000 domains and 216,524 pages across 20 niches, based on around 100,000 prompts, and found referring domains to be the strongest single predictor of ChatGPT citations. Domains with up to 2,500 referring domains averaged 1.6–1.8 citations; those over 350,000 averaged 8.4. Their parallel AI Mode study — 2,328,533 pages from 295,485 domains against 500,000-plus prompts — found the same ladder: 300 referring domains to 24,000-plus moved average citations from 2.5 to 6.8.

So which is it? Both, because they aren't asking the same question. Ahrefs measured brands: does your name appear in the answer? SE Ranking measured domains and pages: does your URL appear in the source list? A publisher-scale link profile predicts being a source; a mention footprint predicts being the recommendation. Treating those as one number is the fastest route to a bad budget decision.

Three statistical tells that this is a threshold, not a slope

If links were a linear driver of AI citations, the data would look linear. It doesn't — and three independent datasets show the same bend.

The step is visible in the raw curve

SE Ranking's referring-domain data isn't a smooth ramp. Below roughly 2,500 referring domains, average citations sit around 1.6–1.8 and barely move as links accumulate. Then, in their words, "the biggest growth spike occurs when crossing the 32,000-link mark, where citations nearly double from 2.9 to 5.6."

Sit with that scale. Thirty-two thousand referring domains isn't a link-building target; it's a media property. At the survey-reported price of $300 to $1,000 per link, buying your way from 2,500 to 32,000 would cost between $8.8 million and $29.5 million. That curve describes what large publishers look like, not a plan you can fund.

Pearson and Spearman disagree — and the gap is the fingerprint

The Semrush and Kevin Indig backlinks-in-AI-search study tracked 1,000 randomly selected domains across ChatGPT, ChatGPT Search, Gemini, Google AI Overviews and Perplexity. The link correlations came in pairs:

Link type Pearson Spearman
Follow links 0.334 0.504
Nofollow links 0.340 0.509
Image links 0.415 0.538
Text links 0.334 0.472

Pearson measures linear association; Spearman measures monotonic rank association. When Spearman substantially exceeds Pearson, the relationship is real but bent — the ordering holds while the straight-line fit doesn't. That's exactly what "clear a bar, and past the bar extra links do little" looks like plotted. Indig got there in words: "Don't expect returns when you just get started or from linear growth. You need a minimum investment to see the expected impact."

Two other results in that table quietly undercut the PageRank story. Nofollow correlated essentially identically to follow, and image links correlated more strongly than text links. If link equity were what these systems responded to, neither should hold. Both are consistent with something simpler: a link is evidence that your brand exists somewhere reputable, and the rel attribute has nothing to do with that.

The Ahrefs sample was already over the bar

Here's the point that makes the 0.664-versus-0.218 comparison more interesting, not less. Every brand in the Ahrefs sample had Domain Rating above 40. That's defensible study design, published openly — but it restricts the sample to the upper part of the link-authority range, and restricting a variable's range mechanically attenuates its correlation with anything else.

So the study isn't answering "do backlinks matter?" It's answering something narrower and more useful: among brands that already have a solid link profile, does adding more links help? Barely. Does adding more mentions help? A great deal. That's the threshold model stated in the shape of the data rather than asserted — and the DR>40 filter sits right on the sweet spot practitioners already target, since 53% of State of Link Building respondents aim at DR 40–60 sites.

One last tell from the same table, and almost nobody quotes it: branded anchors correlated at 0.527, more than twice as strongly as raw backlinks at 0.218. Both metrics count hyperlinks. The only difference is that one carries the brand name as visible text — the cleanest evidence available that the name in the copy does the work while the hyperlink comes along for the ride.

Why the threshold exists: two mechanisms, two budget lines

A threshold with no mechanism behind it is just a pattern. Here there's a plausible mechanism, in two halves.

Half one: retrieval is index-gated, and indexes still respond to links. Seer Interactive ran 100 queries through SearchGPT, extracted 500-plus citations, then searched the identical queries on Google and Bing. Over 87% of SearchGPT's citations matched Bing's top organic results, most in positions 1–10; Google matched only 56%, at a median rank of 17. Being retrievable is a ranking problem, and ranking is still partly a links problem.

But that coupling is loosening fast. In July 2025, Ahrefs found 76.1% of AI Overview citations ranked in the top 10 across 1.9 million citations from 1 million AI Overviews. Seven months later, the same team ran a larger sample against Gemini-3-powered AI Overviews — 863,000 keyword SERPs, 4 million AI Overview URLs — and only 37.9% of cited URLs ranked in the top 10, with 31.0% not ranking in the top 100 at all. Louise Linehan's conclusion: "ranking in the same SERPs as an AI Overview is no longer enough to win an AIO citation." Si Quan Ong had put a floor under it on the earlier data: ranking #1 helps, "but that chance is a coin flip at best."

Part of the reason is that one question is no longer one retrieval. Google's Head of Search, Elizabeth Reid, described AI Mode's query fan-out as breaking a question "into different subtopics" and issuing "a multitude of queries simultaneously on your behalf." Your #1 ranking now competes with the results of a dozen searches you never saw. And 18.2% of those non-ranking citations were YouTube URLs — YouTube is the most-cited domain in that dataset, up 34% in six months, which recasts the 0.737 YouTube correlation as the same finding from the citation side.

Half two: the brand choice may happen before retrieval does. Seer analysed 541,213 LLM responses across 20 brands, 6 platforms and 5 funnel stages, backed by six behavioural tests, and their leading hypothesis is that citations are post-hoc: the model picks which brands to name from parametric memory, then retrieves sources to support the choice. Their phrasing is hard to improve on — "The citations are the bibliography, not the brainstorm." The supporting number is stark: when a brand is mentioned in a response, its citation rate is 53.1%; when that same brand isn't named in the prose, 10.6%. A fivefold gap in one dataset says being named and being cited run on separate machinery.

Put the halves together and the budget structure falls out. Mention-surface presence — your name appearing at scale across many independent sources — determines whether you get named. Index presence and retrieval quality, which links still influence, determine whether you get cited. Two mechanisms, two line items. Funding one and reporting on the other is how teams conclude that nothing works.

So which mentions? The published data is unusually specific, and several of the highest-yield options are cheap.

Earned media does the bulk of the citing. Muck Rack's May 2026 study of more than 25 million links across ChatGPT, Claude and Gemini in 17 industries found earned media accounts for 84% of AI citations, professional journalism alone for 27%, and advertorial or paid content for 0.3%. That 84% has held across all three editions since July 2025, ranging 82–89% — about as stable as anything here gets. The engines use it differently: ChatGPT cites in 96% of responses with ~5 sources (top domain Wikipedia); Gemini in 82% with 8 (Reddit); Claude in just 55% but with 13 (PubMed Central). We've mapped how each engine sources its answers and which tactics measure as zero separately.

Distribution beats domain — and it's the closest thing to causal evidence available. Every correlation above is vulnerable to "big brands just have everything." Stacker and Scrunch ran a study that isn't: eight articles, roughly 189 prompts, 944 prompt-platform combinations across five AI platforms, December 2025. Same article, same brand, same content. On the brand's own domain it was cited 7.6% of the time; distributed through third-party news publishers with canonical tags pointing home, coverage hit 34% — a 325% lift. Stacker sells earned-media distribution, so read it as interested research — though they publish their limits ("the scope was limited to eight stories, five platforms, and citations rather than placement or longevity") and then replicated at larger scale in March 2026: 87 stories, 30 brands, 8 platforms, median citation lift 239%. The effect shrank on replication, which is what honest research looks like.

Review profiles are the cheapest threshold in the dataset. Seer, commissioned by Trustpilot, analysed 804,491 AI responses across 1,926 brands on four engines in March 2026. Brands with no review profile: 1% median citation rate. Brands with as few as 1–13 reviews: 53.5%. A 52-point swing for a claim form and a few review-request emails rather than a $500 link. It skews late-funnel — review and trust sites are 1.51% of citations at awareness but 24.27% at intent — and 99.5% of Trustpilot citations arrived via organic search rankings, meaning the platform's own SEO does the retrieval work for you.

In B2B, the directories consolidated. An analysis of 25,755 AI citations across 200 buyer-intent prompts put G2 at 2.09% citation share on bottom-of-funnel prompts. After G2 acquired Capterra, Software Advice and GetApp from Gartner in February 2026, the group's combined BOFU share rose to 3.68% — a 76% increase, moving it to #2 behind Reddit, and 12.69% on proof queries. A meaningful slice of B2B visibility is now a vendor-relations problem, not a link problem.

Communities and freshness both show dose-response. In SE Ranking's data, domains with heavy Quora presence averaged 7.0 ChatGPT citations against 1.7 for minimal presence, with a similar 1.8-to-7 spread on Reddit — engaging there lets smaller sites "build authority and earn trust from ChatGPT, similar to what larger domains achieve through backlinks and high traffic." The same study found pages updated within three months averaging 6.0 citations against 3.6 for outdated ones. Refreshing a page you already own costs less than a mid-market link and moves a comparable amount.

The best worked example of both halves at once comes from Lily Ray, reported by Search Engine Land. She ran 100 B2B "best [category] software" queries on three dates in 2026 using Ahrefs Brand Radar. Of the 80 that triggered AI Overviews, self-promotional listicles picked up 323 citations — but in 224 of those cases (69%), Google cited a brand's own page while recommending someone else. In one it cited an Oasis LMS article, then recommended four competitors named inside that article. The brands actually recommended shared three traits: category leadership, third-party mentions, and stronger link profiles. Getting cited is not getting recommended — and if commercial queries drive your pipeline, getting into the third-party "best tools" lists engines pull from matters more than another guest post.

One objection worth killing before it reaches your pitch list: "that outlet blocks GPTBot, so the placement is wasted." Citation Labs and BuzzStream analysed 4 million citations from 3,600 prompts across four engines and found 88.2% of GPTBot-blocking news sites still appearing in AI citations. Publisher-side crawler blocking is a far weaker filter on placement value than it sounds.

The budget conversation

Here's the thing about the reallocation this implies: for many teams it's already half-done. In the State of Link Building survey, digital PR was rated best-performing by 34% of respondents, HARO and journalist sourcing by 21%, and guest posts by 18% — PR-shaped approaches taking a clear majority, while link exchanges didn't register at all. The tactics have shifted toward earned coverage. What hasn't shifted is the line item's name and the KPI attached to it. You are probably already buying mentions and reporting them as links.

Then there's the price gap. The survey says 76% of respondents will pay $300 or more per link and 31% pay $500–$1,000. But PressWhizz analysed 22,703 completed placements worth $3.66 million in real transactions and found a median price of $112 and an average of $161. That gap tells you the money is concentrated at the premium end — exactly where link buying and digital PR start costing the same. At $112 you're buying a placement. At $700 you're buying something a PR budget could also buy, and PR returns a named mention in context rather than a hyperlink whose rel attribute the data says barely matters. The same report found links from high-traffic sites selling for roughly 218% more than low-traffic ones — traffic, not Domain Rating, is what the market prices, echoing SE Ranking's finding that domain traffic carried the highest SHAP value of any AI Mode factor tested.

So what's a defensible allocation? Not "stop building links" — below the bar, the threshold model says the opposite. A heuristic that follows from the evidence:

  1. Keep links funded to clear and hold the threshold. Under DR 40, or with key commercial pages not ranking, links are still the constraint. Seer's 87% Bing-match finding says index presence is a prerequisite you can't route around.
  2. Move the top slice of link spend — the $500-plus placements — into digital PR. Same money, different deliverable: a named mention in an earned-media context, where 84% of AI citations come from.
  3. Do the near-free threshold moves first. Claim and populate review profiles. The 1%-to-53.5% step is the highest-ROI item in this entire post.
  4. Fund a distribution line, not just a placement line. The 239% median lift came from putting existing content on other people's domains, not from making new content.
  5. Add video. YouTube mentions correlated at 0.737 and account for 18.2% of non-ranking AI Overview citations. Most link budgets have no line for this.
  6. Refresh before you publish. Updating within three months nearly doubled average citations in SE Ranking's data.

Then keep both sides measured, because none of these ratios are stable — which decides whether the reallocation survives its first quarterly review.

Measuring it so the shift survives the next QBR

The most dangerous way to evaluate a mention-building programme is to count citations. Semrush's ghost citations study with Kevin Indig — 3,981 domain appearances across 115 prompts, 14 countries and 4 engines, published 9 June 2026 — found that 62% of AI citations are "ghost citations": the engine uses the page as a source, but the brand name never appears in the answer. Count citations alone and you will systematically under-report what your PR did.

Worse, the engines fail in opposite directions. Gemini mentions brands 83.7% of the time but cites only 21.4%. ChatGPT is the mirror image: 87% citation rate, 20.7% mention rate. A single blended number averages two opposite behaviours into something that describes neither. Indig's framing of why: "The AI knows the information came from somewhere, but doesn't feel the need to explicitly say so to users. The brand name carries on its own."

Two more findings from that study will wreck any before/after comparison that ignores them. Short conversational queries produce 30 to 50 times more brand mentions than long structured prompts, and comparative queries ("best," "vs") generate 2.4x more than informational ones (43.3% versus 18%). Mention rates by geography ranged from 50% in India and Sweden down to 18–22% in Italy, Brazil and the Netherlands. If your prompt set drifts between readings — a few more comparative prompts, a locale change, a rephrase — you'll measure prompt design and report it as campaign impact. Which is why choosing your tracked prompt set deliberately matters more than most teams assume.

The reporting shape that follows:

  • Mention rate and citation rate, tracked separately. Never blended; the 62% ghost rate is the gap between them.
  • Per engine before averaging. Gemini and ChatGPT are inverses, and a combined figure hides that.
  • Share of answers against named competitors on non-branded buyer-intent prompts — Semrush's "AI share of voice", alongside citation lift and competitive citation displacement.
  • A rolling window with a confidence band, because these systems are nondeterministic. One scan is a snapshot, not a trend.

That reasoning is why Geoptimizer's visibility score weights mention rate at 35% and citation rate at 25%, computes each engine's score before averaging, and reports the headline number as a 7-day rolling window with a confidence band. Because scans run on demand in a minute or two rather than on a daily batch, you can also take a reading the week a campaign lands instead of guessing at the lag.

Be honest about lag in the meeting, too. Semrush's guidance is that "digital PR might take anywhere from days to months to influence AI mentions and citations." Seer are blunter — "none of these changes produce results overnight" — with one case study taking around eight weeks to fully surface. Promise a quarter, not a fortnight.

And budget for volatility. Semrush tracked 230,000-plus prompts over 13 weeks in 2025 across three engines and 100 million-plus citations: ChatGPT's Reddit citation share collapsed from roughly 60% to roughly 10% in mid-September, while Perplexity's barely moved. Anyone who had shifted a quarter of their budget into Reddit-seeding on the earlier number watched it evaporate in weeks. Diversify the mention surface; don't chase one platform.

The objections you'll actually get

"Google says don't chase mentions." It does, and the wording matters. Google's AI optimization guide states: "Seeking inauthentic 'mentions' across the web isn't as helpful as it might seem. Our core ranking systems focus on high-quality content while other systems block spam." Note inauthentic — the same distinction Google draws on links, manufactured versus earned. Nothing there argues against PR, original research, review profiles or product presence.

"Google also says GEO is just SEO." Also true, and quotable: "From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO." Two fair responses. Google speaks for Google, and Ahrefs' engine split (0.65 / 0.30 / 0.15) shows the surfaces don't behave alike. And "it's still SEO" and "your link budget is misallocated inside SEO" aren't in conflict.

"Correlation with what? Big brands have everything." The strongest objection, and every serious source concedes it. Brand size plausibly causes mentions, links, branded search and AI citations at once. Growth Memo matched 7,000 LLM citations to 75,000 brands and found branded search demand the strongest correlate of chatbot mentions at roughly r = 0.33 — arguably evidence that the real driver is demand, with mentions and links both downstream. The best answer available is the Stacker design: hold brand and content constant, vary only distribution, and you still get 325% then 239%.

"So unlinked mentions beat linked ones?" That isn't shown, and it's worth saying so. Nobody has published a study isolating unlinked from linked mentions with the same brand, content and period; Ahrefs' "branded web mentions" metric includes both. The 0.664-versus-0.218 comparison is mentions-in-general against links-in-general. The nearest proxy is the nofollow-versus-follow result — suggestive, not conclusive.

"There's a study saying the opposite." There's a vendor position: a release from the link-building company Link.Build claiming a "strong positive association" between authoritative backlinks and LLM citations, with no sample size and no coefficients disclosed. Treat it as a position rather than evidence; SE Ranking's study makes the pro-links case far better, with published methodology.

"Isn't there a cheap technical fix?" Two popular candidates test poorly: SE Ranking found llms.txt "has shown negligible impact," and pages with FAQ schema averaged 3.6 citations against 4.2 without. Content-level tactics do have peer-reviewed support — the GEO-Bench paper from KDD 2024 found adding quotations, statistics and citations can lift generative-engine visibility by up to 40% — and Backlinko reports pages with expert quotes averaging 4.1 ChatGPT citations against 2.4 without. Content shape is real; a text file in your root directory, on current evidence, is not.

FAQ

Yes, as a qualifying condition. Referring domains were the strongest single predictor of ChatGPT citations in SE Ranking's 129,000-domain study, and 87% of SearchGPT citations matched Bing's top organic results in Seer's test — so links still buy the index presence retrieval depends on. What the data doesn't support is a linear return: Ahrefs found raw backlinks at 0.218 against 0.664 for branded web mentions, and Semrush's Pearson-versus-Spearman gap points to a threshold rather than a slope.

In Semrush's 1,000-domain sample, nofollow links correlated with AI mentions essentially identically to follow links — 0.340 versus 0.334 Pearson, 0.509 versus 0.504 Spearman — and image links correlated more strongly than text links. Engines appear to respond to the mention and its context rather than to link equity, which means a nofollow placement in a strong publication isn't the consolation prize it is in classic link building.

There's no universal split, but the evidence supports a rule: keep enough link investment to clear the threshold and maintain indexation, then reallocate the premium tier of your spend. With a marketplace median link price of $112 and 31% of surveyed teams paying $500–$1,000, placements at the top of that range cost roughly what earned-media work costs — and earned media accounts for 84% of AI citations against 0.3% for advertorial.

How long until digital PR shows up in AI answers?

Longer than a link, shorter than a rebrand. Semrush's honest range is "anywhere from days to months," and Seer note that "none of these changes produce results overnight," with one case study taking about eight weeks to fully surface. Set expectations at a quarter, take a baseline before the campaign lands, and use a rolling window so ordinary nondeterminism doesn't read as movement.

What's the cheapest thing I can do to get cited more?

Claim your review profiles. In Seer's 804,491-response study for Trustpilot, brands with no review profile had a 1% median citation rate; brands with 1–13 reviews sat at 53.5%. Review and trust sites also jump from 1.51% of citations at awareness to 24.27% at intent, so the effect lands where buying decisions happen. After that, refresh existing pages — 95% of ChatGPT citations came from content published in the last 10 months.

Where this leaves the budget

The honest summary is narrower than either headline. Backlinks aren't obsolete, and the correlation studies don't say they are. They say links work as a qualifying condition most established brands have already met, after which additional link spend buys progressively less AI visibility than the same money spent on earned mentions, review profiles, distribution and video. Below the bar, keep building. Above it, the marginal dollar is better spent getting your name into other people's sentences.

What you shouldn't do is decide this from a blog post, including this one. Every ratio here has a scope and a shelf life: 76% of AI Overview citations ranked top-10 in July 2025, and 38% did by January 2026. The only reallocation that survives two quarters is one attached to a measurement you can re-run.

So take the baseline before you move anything. Geoptimizer's free plan covers all four engines — ChatGPT, Gemini, Claude and Grok with no per-engine add-ons, so you can establish mention rate and citation rate separately, per engine, on the prompts your buyers actually ask — then run the same prompts after the campaign and find out whether the reallocation worked. That's a defensible number for the next budget meeting, which is more than 76% of your peers currently have.

Keep reading

See it on your own domain.

Free visibility check across ChatGPT, Gemini, Claude, and Grok — about 30 seconds.

Run the free check