· 20 min read · Geoptimizer Team
Does Wikidata Entity Data Drive AI Search Visibility?
- entity seo
- wikidata
- technical geo
- ai visibility
- structured data
Partly, unevenly, and not in the way most entity-SEO advice claims. The only controlled public test of structured data's effect on AI citations — Ahrefs' difference-in-differences study of 1,885 pages that added JSON-LD between August 2025 and March 2026, matched against 4,000 control pages — found ChatGPT citations up 2.2%, Google AI Mode up 2.4%, and AI Overviews down 4.6%. None of that is a lift. What entity data does instead is decide whether an engine can resolve "you" as a single, unambiguous thing before it decides whether to cite you — and that is an eligibility problem, not a ranking one.
That distinction is the whole post. If you already ship Organization markup, already have sameAs pointing at your LinkedIn and X profiles, and you're wondering whether the next lever is a Wikidata item or a Wikipedia article, the honest answer involves three separate bodies of evidence that mostly don't agree with each other. Let's go through them, starting with the one that should worry you most.
Google deleted three billion entities in a single week
In June 2025, Google ran the largest contraction of its Knowledge Graph in a decade. Kalicube, which has monitored the graph since 2015, recorded a 6.26% drop across two updates on 13 and 20 June — more than three billion entities removed in a week, after a year of steady growth averaging 2.79% a month. The cull was not random. Event entities fell 76.91%. Ambiguous "Thing" entities — the ones Google couldn't confidently type as a company, a person, or a product — fell 15.27%, while unambiguously typed Person entities rose as a share of all persons, from 70.16% to 76.78%. Jason Barnard's summary of what survived: "Clarity is now the only point of entry."
Read that as a directional signal rather than a metric to chase. Google is not loosening its grip on entity identity as AI answers grow; it is tightening it, and the entities it deleted first were the ones it could not confidently classify. If your brand exists in the graph as a vaguely-typed "Thing" with two spellings of its name and no corroborating references, you are in exactly the population that got pruned.
That matters because the Knowledge Graph is not a legacy Search feature. When Google announced AI Mode in March 2025, it described the system as drawing on "high-quality web content, but also tap[ping] into fresh, real-time sources like the Knowledge Graph, info about the real world, and shopping data," and explained the query fan-out technique that issues multiple concurrent searches across subtopics and data sources. The graph is an input to the answer, not an artifact of the blue links.
So the question isn't whether entity resolution happens. It's whether the things you can actually control — your markup, your Wikidata item, your name consistency — are what moves it.
Why identity resolution happens before content quality gets a vote
An AI answer that mentions your brand gets there by one of two routes: the model already knows about you from pre-training, or it retrieves something about you at answer time. Both routes are gated on how well-attested your entity is, and the academic evidence on this is unusually clean.
Kandpal et al. showed that a language model's ability to answer a fact-based question correlates with the number of pre-training documents about the entity involved, and that models up to 176B parameters would need scaling by "many orders of magnitude" to answer questions with little supporting documentation. Retrieval augmentation reduces that dependence, but doesn't remove it. Mallen et al. arrived at the same conclusion from the other side with PopQA, a 14,000-question benchmark: "LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail." The Head-to-Tail benchmark, 18,000 QA pairs split across head, torso and tail facts, concluded that LLMs are "still far from being perfect in terms of their grasp of factual knowledge, especially for facts of torso-to-tail entities" — which is why knowledge graphs and language models are still treated as complementary systems rather than substitutes.
Here is the uncomfortable translation for a mid-market brand: you are the long tail. Not a household name, not in the parametric memory of any frontier model in any reliable way, and therefore answered — when you're answered at all — by retrieval. Your entity work isn't competing with Nike's; it's determining whether a retrieval step can match a page about you to a query about your category.
This is also the concrete reason Wikipedia punches so far above its size in model behavior. In the published GPT-3 training mix, Wikipedia contributed 3 billion tokens — 3% of the mix — but was sampled at 3.4 epochs, while filtered Common Crawl contributed 410 billion tokens at 60% of the mix and just 0.44 epochs. The paper is explicit that datasets "are not sampled in proportion to their size, but rather datasets we view as higher-quality are sampled more frequently." That 3% figure is worth memorizing, because the widely circulated claim that Wikipedia is "around 22% of LLM training data" has no published mix behind it. Use "deliberately over-weighted relative to its size" if you need the shorthand, and skip the number nobody can source.
The Wikipedia number everyone quotes, correctly scoped
Wikipedia's role in AI answers is real, large, and consistently misquoted. Three independent measurements, three different scopes:
| Measurement | Scope | Wikipedia's share |
|---|---|---|
| Profound, 680M citations, Aug 2024–Jun 2025 | All ChatGPT citations | 7.8% |
| Same study | ChatGPT's top-10 sources only | 47.9% |
| Same study | All Google AI Overviews citations | 0.6% |
| Ahrefs Brand Radar, broad US query set, July 2026 | All ChatGPT citations | 8.9% (#2 behind Reddit at 16.7%) |
| Profound, ~730k ChatGPT conversations, Oct–Dec 2025 | Conversations containing citations | ~18% contain a Wikipedia citation (~5% of all citations) |
The 47.9% is the number that gets repeated as "ChatGPT cites Wikipedia half the time." It doesn't mean that. It's Wikipedia's share of ChatGPT's ten most-cited domains — a much smaller denominator than "all citations," where the real figure sits between 5% and 9% depending on who measured and when. If you put the wrong one of those numbers in a client deck, the client will find the correction before you do.
What survives the scoping is still striking. Profound's read of 730,000 conversations calls Wikipedia "the de facto knowledge layer, the place ChatGPT goes first for baseline facts," appearing in roughly one in six conversations that contain citations, with about six unique citations pulled per searching conversation. And the same study found citation concentration is extreme (a Gini coefficient around 0.8) while the top 10 domains capture only 12% of all citations — a long tail with a very heavy head sitting on top of it.
Then there's the engine spread, which is where blanket advice falls apart. Wikipedia is 7.8% of ChatGPT's citations and 0.6% of AI Overviews', and it isn't in Perplexity's top 10 at all. That's more than a tenfold gap between two engines your buyers use interchangeably. Entity-and-encyclopedia work is disproportionately a ChatGPT play, and any measurement that blends four engines into one number will show you roughly a quarter of whatever effect you produced.
Even within one engine, the mix moves for reasons that have nothing to do with you. Semrush tracked 230,000 prompts and over 100 million citations between 14 July and 12 October 2025 and found Wikipedia present in roughly 55% of ChatGPT responses before mid-September, then under 20% after — a platform-level shift isolated to ChatGPT, with Reddit collapsing from around 60% to 10% in the same window. Any before/after test of your own entity work that straddles a swing like that will tell you nothing true.
Where Wikidata actually enters the pipeline
Start with the deflating fact: wikidata.org does not appear anywhere in Ahrefs' top 50 most-cited ChatGPT domains. Not at #43, where LinkedIn sits with 0.9%. Nowhere. If you're expecting a Wikidata item to show up as a citation in your visibility reports, adjust that expectation now — it won't, and a tool that showed it probably measured something else.
Wikidata's value is upstream of the citation layer, and it has become considerably more concrete in the last eighteen months.
It's part of what AI companies license. Wikimedia Enterprise — the paid, structured, high-volume feed — serves Wikidata alongside Wikipedia, Wikivoyage, Wikisource, Wiktionary and the rest, through on-demand, hourly snapshot and realtime APIs. On 15 January 2026, Wikipedia's 25th birthday, Amazon, Meta, Microsoft, Mistral AI and Perplexity joined existing partners including Google and Ecosia. Worth noting for calibration: TechCrunch's report of the announcement observes that neither OpenAI nor Anthropic is named among the partners, so the licensing story doesn't cover every engine you track.
It's now queryable as vectors. The Wikidata Embedding Project, led by Wikimedia Deutschland with Jina.AI and DataStax, launched on 1 October 2025 and exposes the corpus — roughly 119 million items maintained by about 24,000 monthly volunteer editors — through semantic search and Model Context Protocol access, so an agent can query it without writing SPARQL. The stated use cases are RAG grounding, hallucination reduction and fact-checking. Whether any production consumer assistant queries that endpoint today is an open question nobody has answered publicly; treat it as infrastructure that exists, not as a channel with measured traffic.
It's a cross-reference to the graph you actually care about. Wikidata carries a Google Knowledge Graph ID property, P2671, which is also the only place we found a verifiable series for the graph's scale: 570 million records in December 2012, 5 billion in May 2020, 54 billion by March 2024 (an estimate). Numbers like "500 billion facts" circulate widely in secondary SEO writing without a Google primary source behind them — leave those out of your deck.
The practical read: Wikidata is a corroboration layer. It doesn't get cited, it helps other systems agree on who you are. That is a modest, real, and unglamorous benefit — which brings us to the study that says the whole category is oversold.
The uncomfortable evidence: schema alone didn't move citations
Ahrefs took 1,885 pages that added JSON-LD schema between August 2025 and March 2026, matched them against 4,000 control pages, and measured AI citations in ±30-day windows. Four separate statistical approaches — t-test, difference-in-differences, event study and a symmetrical window — agreed: AI Overviews −4.6% relative to controls (statistically significant, roughly 1-in-2,500 odds of arising by chance), AI Mode +2.4% and ChatGPT +2.2%, both indistinguishable from zero. The raw data looked encouraging at first glance, because AI-cited pages were nearly 3× more likely to carry JSON-LD than uncited pages — but the authors demonstrate that correlation is a ranking artifact, since Google's top-10 results are already enriched for schema-bearing pages. Their conclusion: "If you're already doing the rest of the SEO work well, JSON-LD isn't going to be the unlock."
Google says something compatible in its own documentation. The Search Central page on AI features states plainly that there are "no additional requirements to appear in AI Overviews or AI Mode... You don't need to create new machine readable files, AI text files, or markup... There's also no special schema.org structured data that you need to add." The Organization structured data doc goes further: there are no required properties at all — name, alternateName, url, logo and sameAs are all recommended, with logo noted as influencing which logo appears in Search and knowledge panels.
Two honest caveats before you delete your markup.
First, scope. Every page in Ahrefs' sample already had 100+ AI Overview citations before the change. The study answers a specific question — does adding JSON-LD to already-visible pages generate more citations? — and answers it no. It does not test whether a coherent entity identity helps an unrecognized brand get resolved in the first place, which is the actual question a mid-market brand is asking. Nobody has published a controlled test of sameAs, or of creating a Wikidata item, against AI citations. The available material is agency guidance, correlational studies, and one vendor's n=1 self-experiment. That's the state of the evidence, and pretending otherwise is how this category loses credibility.
Second, and more usefully: the strongest measured correlates of AI visibility aren't on your site at all. Across 75,000 brands, Ahrefs found branded web mentions correlating at 0.664 with AI Overview presence, branded anchors 0.527, branded search volume 0.392 — against backlinks at 0.218 and total site pages at 0.17. Brands in the top quartile for web mentions averaged 169 AI Overview mentions; the 50–75% quartile averaged 14. About 26% of studied brands got zero mentions at all. The follow-up across ChatGPT, AI Mode and AI Overviews put YouTube brand mentions at roughly 0.737 — the strongest single signal measured — with site page count at about 0.194. Ahrefs stresses throughout that these are moderate-to-weak correlations and not causal.
That's the reframe worth carrying: mention-earning is the bigger lever, and entity data is what makes those mentions resolve to you rather than to a similarly-named consultancy in another country. It is plumbing, not propulsion. If your name appears in 40 industry articles under three different spellings with no consistent identifier tying them together, you've generated 40 mentions and possibly three entities. That's the failure mode entity work actually prevents — and it's why the correlation numbers and the schema null result aren't in conflict. Which raises the question of whether engines can even see the markup you're relying on.
The crawl-layer problem nobody puts in the checklist
In December 2025, searchVIU built a test page for a fictional product and encoded the same price in eight different formats — visible HTML, JavaScript-rendered, JSON-LD, hidden Microdata, RDFa — then asked five AI systems to read it. The headline finding: "JSON-LD Schema Markup is NOT extracted by ANY system during direct fetch." Gemini found the price in 4 of 8 encodings and was the only system that executed JavaScript. ChatGPT managed 3 of 8, all visible HTML. Claude found it in none. Ahrefs' schema study reports the same behavior: during direct retrieval, every system extracted only visible HTML, ignoring JSON-LD, hidden Microdata and hidden RDFa alike.
If you build entity pages the way most technical SEOs do — clean visible copy plus a comprehensive JSON-LD block carrying the founding year, headquarters, legal name and identifiers — the engines are reading half your work. The half they're reading is the copy.
There's a sharper version of this failure that catches modern stacks specifically. JSON-LD injected client-side, through a tag manager or a streaming SSR framework, is fine for Googlebot and invisible to crawlers that don't execute JavaScript; developer discussion of the streamed-metadata case in Next.js is a useful primer if your stack is React-based. Fetch your own page without JS and read what comes back. If the founding year, HQ city, category description and product names aren't in that HTML, no amount of sameAs will help, because nothing is reading the node they live in. We walk through the wider version of this test in our post on llms.txt and AI crawler access, where readable no-JS HTML is one of the four twenty-minute checks.
Access is the precondition underneath all of it. OpenAI documents three separate agents, and only one governs whether you can appear in ChatGPT's search answers: GPTBot for training, ChatGPT-User for user-triggered fetches, and OAI-SearchBot, which surfaces sites in ChatGPT's search features — sites blocking it "will not be shown in ChatGPT search answers." Blocking rules have a habit of drifting, especially at the CDN layer; our bot-by-bot guide to the Cloudflare crawler rules covers how to keep retrieval crawlers in while making a deliberate choice about training ones. Geoptimizer's free AI Crawler Checker and GEO Site Audit run those checks — crawler access, structured data, content shape — without a signup, which is the fastest way to rule out a plumbing problem before you invest a quarter in entity work.
A defensible entity checklist, with the odds stated
Nothing below is a guaranteed citation lift; every item is either an identity assertion or a removal of a known blocker. That's the honest framing.
- Fix one canonical identity. One legal name, one brand name, one
@idURI used on every page. Put the brand name innameand the legal name inalternateName. This is cheap and it's the thing the June 2025 Knowledge Graph cull rewarded. - Treat
sameAsas identity, not links. Schema.org defines it as "URL of a reference Web page that unambiguously indicates the item's identity" — Wikipedia, Wikidata, the official website. Google's Organization doc reads it more loosely as profile pages elsewhere. Add profiles you actually control and that name you consistently. Do not add it for link equity: external links on all Wikimedia wikis have beenrel="nofollow"since the English mainspace exemption was removed on 20 January 2007. - Duplicate every entity fact in visible HTML. Founding year, HQ, category, product names, leadership. If it exists only in JSON-LD, the direct-fetch crawlers never see it.
- Verify the markup survives with JavaScript disabled. See above; this catches GTM injection and streamed metadata.
- Confirm OAI-SearchBot and the other retrieval agents get a 200. Robots.txt, CDN rules, WAF challenge pages.
- Look up whether Google already has an entity for you. The Knowledge Graph Search API still returns
kg:/m/…IDs, though Google labels it legacy, is migrating it to Cloud Enterprise Knowledge Graph, and warns it is "not suitable for use as a production-critical service." Record what you find as P2671 on your Wikidata item. - Create or clean a Wikidata item. The bar is much lower than Wikipedia's: an item qualifies if it describes "an instance of a clearly identifiable conceptual or material entity that can be described using serious and publicly available references." The company property set worth populating is instance of (P31), official website (P856), inception (P571), headquarters (P159), country (P17), legal form (P1454), industry (P452), plus identifiers — LEI (P1278), ISIN (P946), X username (P2002), LinkedIn company ID (P4264), Crunchbase ID (P2088), CEO (P169), parent organization (P749). Source every statement; a promotional item with no serious references invites a deletion request.
- Claim your knowledge panel if one exists. Google says panel data comes from many sources, including "verified entities who have suggested edits to facts on their own knowledge panels."
- Treat Wikipedia as a consequence, not a task. WP:NCORP requires significant coverage in multiple independent secondary sources and explicitly excludes press releases, sponsored articles, corporate filings, and routine coverage of funding rounds, hiring and product launches — "a single significant independent source is almost never sufficient." If you or an agency edit on a client's behalf, the Wikimedia Terms of Use require disclosing "each and any employer, client, intended beneficiary and affiliation."
For a worked example of the realistic end state, look at Wikidata item Q107533769 — Ahrefs the company. It carries instance-of values for brand, enterprise, web service and SaaS; official website; inception 2010; headquarters and country; official legal name; industry values; a Crunchbase ID; an X handle; and a Google Knowledge Graph ID. Its only Wikipedia sitelinks are Bulgarian, Portuguese and Ukrainian — there is no English article. That is the achievable pattern for a mid-market brand: a complete, sourced, unambiguous Wikidata item and no Wikipedia article at all.
Context for perspective: Web Data Commons' extraction from the October 2024 Common Crawl found structured data on 16,525,070 of 37,447,141 domains — about 44%, and 74 billion triples. Markup is table stakes. Coherent identity across it is not.
How to tell whether it worked
Entity work is slow and its effects are small, which makes it exactly the kind of change that gets falsely confirmed by a badly designed measurement. Four rules keep you honest.
Measure per engine, never blended. A tenfold gap between ChatGPT and AI Overviews means an entity-driven gain that lands on one engine gets diluted to near-invisibility in a four-engine average. This is why Geoptimizer computes a score for ChatGPT, Gemini, Claude and Grok separately before averaging them — the per-engine breakdown is where a ChatGPT-weighted change like this actually shows up.
Segment by prompt type. Magna's vendor study of 5,127 prompts (on ChatGPT-4o, Claude 3.5 Sonnet and Gemini 1.5 Pro) used semantic similarity to detect Wikipedia's influence even where it wasn't cited, and found 58% influence on definitional queries against 14% on comparisons, 11% on local/service and 8% on recommendations. Single vendor, single run — treat it as a hypothesis to test, not a fact. But if it holds, encyclopedic entity work helps most on "what is [brand]" and least on the "best X for Y" prompts your pipeline depends on, and blending those prompt types together will average your effect to zero.
Use a rolling window, not a screenshot. AI answers are nondeterministic. A single scan is a snapshot; a headline number should be a rolling window with a confidence band, or you'll read sampling variance as progress.
Don't run the test across a frontier model release. Semrush's data shows ChatGPT's Wikipedia share moving from ~55% to under 20% inside one mid-September window for reasons unrelated to any brand's entity work. If your before/after straddles a release, you've measured the release. Our guide to separating a model-version artefact from a real visibility change covers the diagnostic in detail.
One more piece of tooling context, since it shapes what you can even verify: Microsoft retired the Bing Search APIs on 11 August 2025, pushing users to Grounding with Bing Search inside Azure AI Agents, and Google's Knowledge Graph Search API is on legacy status. The entity-lookup tooling technical SEOs relied on is quietly disappearing, which makes measuring outcomes in the answers themselves more important than measuring proxies.
FAQ
Do I need a Wikipedia article to get cited by AI? No. Wikipedia is a heavily-used retrieval source — 8.9% of ChatGPT citations in Ahrefs' July 2026 data — but that's Wikipedia being cited, not you. An article about your company is realistic only if independent secondary sources have already covered you in depth, since NCORP explicitly rules out press releases and routine funding, hiring and launch coverage. Earn the coverage; the article is a downstream consequence.
Will a Wikidata item show up in my AI visibility reports?
Almost certainly not. wikidata.org doesn't appear in the top 50 domains ChatGPT cites. Its value is upstream — corroboration for knowledge graphs, licensed feeds and RAG pipelines — so judge it by whether engines describe your brand accurately, not by whether the item gets cited.
Does sameAs pass SEO value?
Not as link equity. Every Wikimedia project tags external links rel="nofollow", and schema.org defines sameAs as an identity assertion — a statement that two URLs refer to the same thing. Adding profiles you don't control or that name you inconsistently makes identity resolution harder, not easier.
Is JSON-LD still worth adding if AI crawlers ignore it on direct fetch? Yes, for classic Search features and for engines that read your pages after indexing rather than by direct fetch — just don't expect citations from it, and don't let it be the only place your entity facts live. Server-render it, and mirror the same facts in visible copy.
How long does entity work take to show up? Nobody has published propagation timings from a Wikidata edit into Google's Knowledge Graph, so treat any specific promise with suspicion. Plan for quarters, measure with a control group and a rolling window, and expect a small effect concentrated on definitional prompts and on ChatGPT.
The honest verdict
Entity data is an eligibility and accuracy lever with strong mechanistic support and weak causal evidence for citation lift. The mechanism is well-documented: models are worse on long-tail entities, retrieval is what answers questions about brands like yours, Google's AI surfaces read the Knowledge Graph, and that graph just demonstrated it will delete anything ambiguous. The causal evidence is thin: the one controlled test found no lift from schema, and no one has tested sameAs or Wikidata against citations at all.
So do the work, cheaply, and price it correctly. Fix the canonical name, get the facts into visible HTML, keep the retrieval crawlers unblocked, build a properly sourced Wikidata item, claim the knowledge panel — then spend the rest of the budget on earning the mentions that correlate at 0.66, because entity data is what makes those mentions resolve to you rather than to someone else with a similar name.
And measure it per engine, over a rolling window, with the model-release calendar open next to you. You can run a free AI Visibility Check across ChatGPT, Gemini, Claude and Grok to see which engines already resolve your brand correctly — no signup, and a first score takes about 90 seconds — or compare what each Geoptimizer plan tracks if you want the per-engine breakdown on an ongoing schedule. Knowing which engine your entity work landed on is the difference between a result and a story.