· Updated · 16 min read · Geoptimizer Team

AI Visibility by Country and Language: Measure Per Market

  • generative-engine-optimization
  • ai-visibility
  • measurement
  • international
  • prompt-sets
AI Visibility by Country and Language: Measure Per Market

If you report one AI visibility number to your exec team, it is almost certainly an English-language number wearing a global label. In a controlled study of more than 16,000 AI answers, Weglot found that Japanese-language queries sent 26% of citations to .jp domains versus roughly 1% for the identical questions asked in English — and French-language queries sent 16% to .fr versus 1.4%. Language, not geography, is the variable that rewires which sources an engine pulls in. Which means a global score is not an average of your markets; it is a measurement of one market, extrapolated. The fix is unglamorous but cheap: a separate prompt set, in the local language, per market you actually sell in.

The average that describes nothing

Start with the arithmetic problem, because it is the one that survives contact with a board meeting. As Echowi puts it in its guide to share of voice in AI answers: "A brand that holds 40% share of voice in one market and 0% in three others reports 10% when you average it. That number describes nothing that happened anywhere, and it hides the only fact worth acting on."

That is the shape of most international AI visibility reporting today — not because anyone is being sloppy, but because the default configuration of almost every measurement setup is English prompts. Your own, your agency's, and the industry benchmarks you quote in the deck. The largest public benchmark in the category makes it explicit: Semrush's expanded 2026 AI Visibility Index analysed 126 million U.S. AI search prompts from January to April 2026 across 22 industries. Excellent work, and a US read. When a growth lead says "we're at 34 against a category average of 41," the category average is usually American.

Meanwhile the users have already moved. OpenAI's June 2026 Signals data, reported by Search Engine Journal, shows more than half of ChatGPT's consumer users now predominantly use a language other than English, led by Spanish, Portuguese and Arabic, with the fastest relative growth in Africa and Asia. If your DACH region is 22% of ARR and your measurement is 100% English, you are not measuring 22% of your revenue. You are guessing about it, monthly, with a decimal point attached.

So how different can it really be? Different enough that the answers are not comparable.

Language rewires the citation graph

The strongest evidence here is recent, large and independent. Profound analysed 3.25 billion AI citations across seven models and 14 countries in March 2026, filtering every prompt to the country's native language. Their conclusion: "The language of a query can rewire the entire citation graph: which domains appear, how often they're cited, and even whether 'the social web' shows up at all."

Two details from that dataset are worth pinning to your wall. First, same-language markets behave alike even on different continents: in Google AI Overviews, Mexico's social citation rate was 22.5% and Spain's 20.0%; in ChatGPT the same pair sat at 4.70% and 4.24%. Geography is secondary to language. Second, the reason ChatGPT's social layer thins out abroad is structural — Reddit accounted for between 51% (UAE) and 76% (India) of ChatGPT's social citations in every country measured. Reddit is English-first, so an engine that leans on Reddit inherits an English-first view of your category no matter where the user is sitting.

That is what happens to sources. What happens to brands is sharper still. In a study published in June 2026, Dmitrij Żatuchin tested 66 brands across 11 Northern, Baltic and Central European markets in 12 languages, collecting 35,640 responses from three grounded models (GPT-5.4, Gemini 3.1 Pro and Perplexity Sonar Pro). Switching a buyer-intent query from English to the brand's home language raised recommendation share by 0.80 for local champions but only 0.15 for global multinationals on a 0–1 scale. Read that from the seat you actually occupy: if you are the challenger in Poland or Finland, an English-only audit systematically understates you and flatters the multinational you are trying to displace. The paper says it plainly — "An English-only audit therefore understates a locally headquartered brand's AI visibility, while representing a multinational's fairly."

Now the nuance that keeps this honest, because it is the strongest counter-argument available and it comes from the same paper. How the models describe brands barely moves across languages: mean cross-language cosine similarity was 0.825, and model choice explained far more response variance than language did (eta-squared 0.32 versus 0.01). Żatuchin's summary: "Query language changes which brands the model recommends far more than how it describes them."

The practical translation is a rule you can apply today. If you are monitoring reputation and sentiment, English-only monitoring is a defensible approximation. If you are monitoring whether you make the shortlist — recommendation share, mention rate, the metric you are actually compensated on — English-only monitoring is measuring somebody else's market.

Why it happens: your answer was researched in English

None of this is a model-fluency problem, and it is worth saying so before someone in the room objects that "the models speak German fine now." They do. Published multilingual benchmarks show parity closing fast, with English retaining only a modest reasoning edge. The gap is in retrieval — which documents get pulled into the context window — not in whether the model can write a clean German sentence.

Peec AI measured exactly that layer, analysing over 10 million prompts and 20 million query fan-outs — the background searches ChatGPT runs before it answers. For non-English prompts, 43% of those fan-out queries ran in English anyway, and in nearly 78% of cases the model decided native-language sources alone were not enough. Turkish topped the list at 94%; Spanish was lowest at 66%; no language fell below 60%. Your German prospect asks in German, and more than four in ten of the background searches behind the answer run on the English-speaking web.

The failure modes this produces are concrete rather than theoretical. Asked in Polish, from Poland, about the best auction portals, ChatGPT either omitted Allegro.pl — the country's dominant marketplace — or buried it beneath global competitors. A German-language query about German software companies returned excellent companies, none of them German. And in the most reproducible demo in the whole dataset, a Spanish query about cosmetics brands had the word "globales" quietly added to the background search, which is precisely how locally dominant brands fall out of an answer that was, on its face, asked in their own language.

Weglot's controlled test explains the other half of the mechanism: ask in a language and the engines favour sources in that language, from that place, but they also pull from a thinner pool — "English queries pulled more sources than local-language queries in every single market we tested." Fewer slots, more local competition for them. Their blunt version: "Content living only in English sits largely outside the pool these systems draw on when they answer in a local language."

There is a delivery-layer wrinkle too. In hands-on tests across nine engines, Glenn Gabe found ChatGPT and Claude frequently returning English URLs despite non-English settings, while Bing, Copilot and Google's own surfaces handled language variants correctly. That work is from December 2025 and the engines have shipped many versions since, so treat it as a hypothesis to re-test against your own hreflang setup — but if your translated pages exist and never surface, it is a plausible reason why.

Your engine mix is not global either

Even if you fixed the language problem, there is a second geographic assumption hiding inside any blended score: the engines you average.

SE Ranking's analysis of Google Analytics data from 101,574 websites between January 2025 and April 2026 found the AI referral mix varies sharply by region. In the US: ChatGPT 71.33%, Gemini 12.87%, Perplexity 6.85%, Copilot 5.12%, Claude 3.70%. In the EU: ChatGPT 75.59%, Gemini 9.8%, Perplexity 9.02%. In the UK: ChatGPT 85.28%, Gemini 5.81%, Perplexity 3.06%. The US is the most distributed mix of any region analysed; the UK is still mostly a ChatGPT market. So a four-engine average weights Gemini more than a UK-only business should care about, and under-weights Perplexity for an EU one.

Country-level idiosyncrasies compound it. Mistral takes 0.85% of AI-driven referral traffic in France versus 0.24% across the EU — roughly 3.5× better in its home market, and an engine that simply does not appear in a US-designed tracking stack. Across APAC, Otterly's per-country review puts ChatGPT first everywhere except China, where Baidu's Ernie Bot leads, with local players like Naver's HyperCLOVA X mattering in South Korea.

The honest response is not to pretend a tool covers everything. It is to read the per-engine rows rather than the combined number whenever a specific market is the question, and to say out loud which engines your measurement misses there. Geoptimizer runs every active prompt against ChatGPT, Gemini, Claude and Grok on every plan, and publishes the evidence behind its 0–100 scoring formula — including that the headline number is the mean of the per-engine scores. Knowing what the average is made of is the point. It also does not cover Ernie Bot or HyperCLOVA X, which is a sentence your China and Korea plans should contain.

The surface itself isn't the same everywhere

There is a third assumption underneath the first two: that the thing you are measuring existed in that market last quarter.

Where answers exist but generate no clickthroughs, you'll need traffic‑independent measurement techniques — see Zero-Click AI Answers: Measuring Visibility Without Traffic for practical methods to detect and quantify answer‑layer presence without referrals.

Google switched on AI Overviews and AI Mode for French searchers on 22 July 2026; nine other European markets — Germany, Italy, Spain, Poland and others — had AI Overviews from 26 March 2025. Google France's managing director attributed the delay to regulatory hurdles "which we believe we have now overcome." So a French AI-visibility trend line that reaches back before July predates the product existing there, while the German one does not. That is not a trend; that is two different products on one chart.

The same applies to language coverage. Google's February 2026 AI Mode expansion added 53 languages in a single release, taking the cumulative total toward 100 and including Czech, Danish, Estonian, Finnish, Greek, Hungarian, Polish, Romanian, Swedish, Arabic and Hebrew. Markets that scored zero before that date scored zero because the surface was not there.

And where the surface does exist, its density varies. Using a dataset of 108 million AI Overview queries, SEOProfy reports coverage rates of 37.2% in Indonesia and 29.1% in Mexico against 20.5% in the United States and 19.1% in the UK. Identical effort in two markets meets a different amount of AI answering.

Even inside one country and one language, location moves the citations. SE Ranking tested 100,013 keywords across five US cities: AI Overview appearance rates differed by less than a percentage point, but only 47.05% of queries cited identical sources across all five locations, and 6.34% had zero domain overlap — concentrated in Legal, Healthcare and Real Estate. If Denver and Los Angeles can disagree completely, Madrid and Mexico City certainly can.

What a market-by-market prompt set actually looks like

Here is the part that turns all of the above into a week of work rather than a headcount.

Pick two to four markets, not twelve. BotSee's scoping advice is to choose markets where revenue already exists, expansion is planned, competitors dominate AI answers, or leadership needs localization evidence — and to build a shared baseline of 10–15 high-intent questions used across all markets, then add 10–20 market-specific variants per country. Their governing principle for the local versions: "Use the same intent structure, not the same wording."

Build matched prompt families, not translations. Answer Engine Land's method uses three prompt types per priority topic: a semantic match (how a local buyer would natively ask it), a literal control (a close translation, kept specifically to isolate translation-driven differences), and a local-intent variant (natural phrasing plus city, country, currency or eligibility qualifiers). Native reviewers validate the wording, because "machine translation can be useful in drafting, but it should not be treated as proof that two prompts have equivalent intent." Source local phrasing from sales-call notes and local competitor research instead.

Log language and location as separate variables. They are separate experimental treatments: "A German-language prompt from Berlin tests something different from an English prompt that asks for options in Berlin." For every observation, record the engine and surface, the prompt language, the declared user location, and the exact wording — plus date, sign-in state and device class. Language sets the source pool; location sets the local detail.

Keep the outcomes separate. Answer presence, citation presence, brand accuracy and source selection are four different findings, and collapsing them into one "visibility" figure is listed explicitly as a common mistake. A market where you are mentioned but never cited needs different work from one where you are cited but described with last year's pricing.

If your brand is a common word, expect AI mention counts to inflate with unrelated uses — see how common-word brand names cause AI mention false positives for practical detection and mitigation techniques.

Then read the gaps diagnostically. This table is the most usable artefact in the literature:

Pattern observed Likely explanation
Different answer after literal translation only Wording or translation effect
Different answer only when city or country is named Local intent or availability
Same answer, different citations Source-ecosystem variation
Correct global answer, wrong local detail Entity or policy ambiguity
Unstable citations in every market Sampling variation or product change

On budget: two markets at roughly 10 shared plus 10 local prompts each lands around 30 tracked prompts — what Geoptimizer's Pro plan tracks at $49 a month, with all four engines included rather than sold as per-engine add-ons. Because pausing a prompt frees its slot immediately, you can also rotate market sets instead of buying a bigger allowance. If you have never built a prompt set this way, our walkthrough on choosing the 25 prompts you track for AI visibility covers the funnel allocation this method assumes as its shared baseline.

One caveat before you trust a location setting. Both OpenAI's web search tool and Perplexity's API accept an approximate user_location with an ISO country code, but that is a hint to the retrieval layer, not proof the engine behaved like a local user. Do not imply that a browser language, VPN or location parameter guarantees an engine used that geography.

How not to fool yourself

Splitting one number into six creates a new hazard: six numbers, each built on a sixth of the sample.

The sampling maths is unforgiving. At a true 30% share of voice and 95% confidence, 10 runs per prompt gives you ±28 points, 30 runs ±16, 100 runs ±9 and 400 runs ±4.5. Reliably detecting a move from 30% to 35% takes roughly 1,400 runs per condition — and every market × language × engine you add is another condition. So when Spain shows 41 and Italy shows 29 off a handful of runs each, you have not found a market gap. You have found a confidence interval.

Three habits keep this honest. Report counts alongside proportions — "cited in 7 of 12 runs" is harder to over-read than "58%". Freeze the prompt set, because a changed prompt set invalidates every comparison against history. And let a chart say "inconclusive" for a small market rather than publishing a tidy country ranking built on two observations. This is the same signal-versus-noise discipline applied to a new axis; if you have read our diagnostic for telling a real visibility drop from model-version noise, you already know the moves — a rolling window, a confidence band, and no rewrites during a release week.

Finally, reset expectations about what a "good" market score looks like. In Weglot's SaaS category, HubSpot — a category leader by any measure — took under 3% of citations. Being cited is less about unseating one dominant name than about being in the mix at all.

What to put in front of your exec team

Replace the single number with a table: market, language, engine, mention rate, citation rate, and the run count behind each row. It fits on one slide and it survives questions, which the single number does not.

If a headline figure is politically required — and it usually is — make it a revenue-weighted composite, with the market rows always visible underneath it. A weighted number at least describes your business. An unweighted average of four markets describes a business that does not exist.

Then expect the rows to disagree with each other, and treat that as the finding rather than a data-quality problem. English content bleeds into local answers far more readily than local content bleeds into English ones, which means your strongest market on paper may be the one where you did the least work, and your weakest may be a market where a local champion holds the shortlist. That is a budget conversation, not a dashboard bug.

FAQ

Should I segment by language or by country? By language first, if you can only pick one. Profound's 3.25-billion-citation analysis found same-language markets pattern together across continents — Mexico's and Spain's social citation rates sat within 2.5 points of each other in Google AI Overviews. But log both: location still moves citations with language held constant, as the five-city US test showed with 6.34% of queries sharing no cited domains at all.

Can I just translate my English prompt set? Use translation for drafting and keep a literal translation as a control, but do not treat it as your primary market prompt set. Machine translation is not proof that two prompts carry equivalent buyer intent, and the differences you care about — a Spanish query silently gaining the qualifier "globales" in its background search — show up exactly at the wording level.

How many markets can I realistically track? Two to four. A programme of roughly 10 shared plus 10 local prompts per market keeps two markets inside a 30-prompt allowance, and pausing prompts frees slots immediately if you want to rotate. Adding markets faster than you can run each condition enough times produces confident numbers with no statistical basis underneath them.

The models handle our languages well now — isn't this solved? Fluency is largely solved; retrieval is not. The measured gap is in which documents get pulled into an answer, not in how well the model writes the language. That is why 43% of ChatGPT's background searches for non-English prompts still run in English.

If AI describes our brand the same way in every language, why bother? Because description is not recommendation. Cross-language response similarity measured 0.825, but recommendation share moved by 0.80 for local champions when the query switched to their home language. Sentiment travels; shortlists do not — and shortlists are what your pipeline is made of.

Measure per market, per language, per engine

A global AI visibility score is not wrong so much as it is a category error: it answers a question about one language as if it answered a question about your whole business. The evidence from three independent large-N studies points the same way — language reshapes the citation graph, local-language queries favour local sources, and an English-only audit quietly flatters your multinational competitors.

You do not need a new team to fix it. You need two or three markets, a native-phrased prompt family for each, run counts attached to every claim, and a report that shows the rows rather than the mean. Start with a free AI Visibility Check and the other free GEO tools on Geoptimizer — run one buyer question in English and again in your local language, and see whether the same brands come back. That five-minute test usually settles the argument before the budget conversation starts.

Keep reading

See it on your own domain.

Free visibility check across ChatGPT, Gemini, Claude, and Grok — about 30 seconds.

Run the free check