Skip to main content

folkfox

Skip to main content
Skip to content
SEO AND GEO

The English page nobody reads is the one the machine reads

Assistants do their reading in English, whatever language you typed. Multilingual SEO now has to serve a machine's research step, not just a human's browse, and the English page nobody visits is suddenly the one that matters.

Quick answerMultilingual SEO now has an odd new rule. AI assistants research in English whatever language the user typed, so one genuinely written English version of your highest-intent pages earns citations your local folders cannot.
Section 01

Two datasets, one uncomfortable finding#

multilingual seo

A fox does not read the whole wood. It reads the one hedgerow where the wind carries something worth chasing, and ignores the rest with a confidence that looks almost rude. Two datasets now sitting side by side hand multilingual SEO the same awkward habit, because the hedgerow the machines keep returning to is written in English.

The first comes from Peec AI, whose analysis covered over 10 million ChatGPT searches and over 20 million query fan-outs from recent months. The data was filtered so that user location matched query language, Polish from Poland, German from Germany, with mismatches excluded. On that filtered set, 43% of research steps were conducted on the English-speaking web, even when the original question was asked in another language.

The per-language picture is sharper still. In Peec AI's sample, Turkish prompts switch to English 94% of the time and Spanish prompts 66%, and no non-English language falls below 60%. In nearly 78% of cases, ChatGPT determines that native-language sources are not enough and goes looking elsewhere.

0%

of ChatGPT research steps ran on the English-speaking web, even when the prompt was written in another language

Peec AI, over 10 million searches

How often the research switches to English
Bar chart of English research rates: Turkish 94 percent, Spanish 66 percent, stated floor for any non-English language 60 percent, and 43 percent of all research stepsTurkish: 94Spanish: 66Lowest floor: 60All research: 4394%70.5%47%23.5%0%TurkishSpanishLowest floorAll research
The single takeaway for multilingual SEO: whatever language the prompt arrives in, much of the research behind the answer happens in English. Turkish and Spanish are the only two languages Peec AI names; 60% is the stated floor no non-English language falls below, not a language; 43% is the share of all research steps run on the English-speaking web.

Hold that against the shape of the web itself. The same piece notes that 80% of internet users do not speak English as a first language while 50% of internet content is written in English, a split that squares with W3Techs, which puts English at 49.5% of all websites whose content language is known, ahead of Spanish and German at 6.0% each. The model is not being snobbish, it is foraging where the food is, and English is the thicket with the most cover.

One honest caveat before anyone spends on this. The Peec AI piece gives no precise collection window beyond "recent months", and does not detail how individual fan-outs were classified as English versus native-language research. That gap is real. It does not sink the finding, but it makes 43% a direction of travel rather than a fixed constant, and multilingual SEO plans should be built on directions.

Nor does it mean your readers have quietly become English speakers. Eurostat found 74.7% of EU working-age adults knew at least one foreign language in 2022, but only 27.6% of them called themselves proficient in their strongest one. The humans still want their own language. Only the research step has changed sides, and that single distinction is the whole of multilingual SEO right now.

Section 02

The research step is not the reading step#

Multilingual SEO only makes sense now if you separate two things that used to be one event. A person searched, a person clicked, a person read. Now a model searches on their behalf, reads a dozen pages nobody will ever see, and hands back a paragraph. Your page can win the research and lose the reading, or win neither, and ordinary analytics will report both outcomes as silence.

Google names the mechanism itself. In its AI optimisation guide, query fan-out is defined as a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results. Ask how to fix a lawn full of weeds and the system quietly asks itself about herbicides too.

How a fan-out spends its attentionIllustrative, not measured. One prompt becomes several concurrent searches, and on the evidence above a large share of them go looking in English, so your English page competes at the research step rather than the reading step.One Dutch promptEnglish how-toDutch reviewEnglish pricingEnglish specsEnglish how-toDutch reviewEnglish pricingEnglish specs
Illustrative, not measured. One prompt becomes several concurrent searches, and on the evidence above a large share of them go looking in English, so your English page competes at the research step rather than the reading step.

That reframing changes what a page is for. A brochure page is written for a browsing human with a mouse and a mood. A research page is written for a system that will lift one sentence, check it against two others, and discard anything it cannot verify alone. The same words rarely do both jobs well, and most brands have only ever laid the first scent.

A page that assumes a human is scrolling

Hero image, a warm brand promise, three benefit tiles, pricing behind a click, and the licence, specification or delivery detail sitting four scrolls down inside a collapsed accordion. Lovely to look at. Almost nothing in it survives being lifted out as a single standalone sentence.

A page that assumes a machine is fetching

One direct answer inside the first forty words, prices and numbers in plain text near the top, every claim self-contained and sourced at the point it is made, headings phrased as the questions people actually ask, and no fact that only makes sense after the paragraph above it.

Peer-reviewable work points the same way. The preprint What Gets Cited: Competitive GEO in AI Answer Engines ran 252,000 trials across six models, testing 18 content factors in paired comparisons. Topical relevance and list position drove being cited first. Explicit pricing data and recent timestamps helped. Formatting alone did almost nothing, which should retire a lot of tidy advice about bullet points.

Know which crawler your multilingual SEO is courting#

Access is the unglamorous half of multilingual SEO, and it is where a lot of English folders quietly fail. OpenAI's crawler documentation separates the jobs plainly: OAI-SearchBot is used to surface websites in search results in ChatGPT's search features, while GPTBot crawls content that may be used in training the foundation models. They are not the same decision, and blocking one to make a point about the other is how brands lose visibility they meant to keep.

Shut out OAI-SearchBot in your robots file and no amount of careful English writing reaches the research step. The gate has to be open before the trail is worth laying, and that four-minute check beats a quarter of content strategy.

Section 03

Multilingual SEO's new lift is real, the samples are small#

The second dataset is the one that turns an observation into a decision. Writing in Search Engine Land, Pieter Serraris measured what actually happens to sites that carry an English folder alongside their local ones. The panels differ per engine: 272 Bing properties for Copilot, 31 of them with full /en/ splits; 26 multilingual properties for ChatGPT, with a five-site subpanel over a 30-day window; and a Search Console panel of multilingual Belgian and European sites for Google AI.

On those panels, ChatGPT's retrieval fetches English pages 65% to 79% of the time. Copilot sites with an English folder saw 892 citations per 10,000 impressions against 585 without, a 52% aggregate difference. Google AI showed a 28% aggregate lift in AI impressions with an English folder, and a median per-site lift of 9% once outliers were removed. ChatGPT showed a 122% uplift.

The English-folder lift, by engine
ChatGPT citations
+122%
Copilot citations
+52%
Google AI impressions
+28%
Google AI, median site
+9%
Read the panel behind each bar before you read the bar. ChatGPT's +122% rests on 26 multilingual properties with a five-site subpanel over 30 days, easily the thinnest number here; Copilot's +52% comes from 272 Bing properties, 31 with full /en/ splits; Google AI's +28% aggregate and +9% median per-site come from a Search Console panel of multilingual Belgian and European sites. All figures from Search Engine Land.

The same piece reports an English-preference index per engine: roughly 2.6 for ChatGPT, 1.07 for Copilot, 0.79 for Google AI. That is an index against a neutral 1.0, not a percentage, and it should never be charted as one. Read plainly, ChatGPT reaches for English far more than a neutral system would, Copilot barely leans, and Google AI actually sits below neutral, which is to say it shows no English preference at all.

rely on small samples ranging from single digits to the low double digits of sites, so treat those results as indicative
Pieter Serraris, Search Engine Land
In nearly 78% of cases, ChatGPT determines that native-language sources aren't enough
Tomek Rudzki, Peec AI
Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users
Google Search Central, spam policies

Put the three quotes side by side and the sensible response appears in outline. Two independent panels agree on a direction. Neither is big enough to bet a rebuild on. The third quote is the wall multilingual SEO hits if it overreacts, which is exactly what most sites will do.

Section 04

Where this becomes scaled content abuse#

Here is the counterweight, from the only party that can penalise you for getting it wrong. Google's AI optimisation guide states that SEO best practices continue to be relevant because its generative AI features are rooted in core Search ranking and quality systems. There is no separate discipline to buy, only a search experience, and this is still SEO wearing a newer hat.

That guide then closes the door on the lazy version of today's finding. Building multiple pages per query variation, it says, primarily to manipulate rankings or generative AI responses, violates Google's scaled content abuse spam policy, and a high quantity of pages does not make a website higher quality or more relevant to users. Nobody gets to read the 122% number and generate four hundred English near-duplicates.

The spam policies page is blunter again. Its named examples include using generative AI tools to produce many pages without adding value, and scraping content to generate many pages through automated transformations like synonymising, translating, or other obfuscation techniques where little value is provided. Machine translation at volume is named in the policy. Not implied, named.

One page, written. Not forty, generated#

So the defensible version of multilingual SEO best practices is narrow and slightly boring. Choose the handful of pages carrying real commercial intent. Write the English version from scratch, with a person who knows the subject. Let it say something the local page does not, because the reader it is written for is a research step, not a browser. Then stop, and leave the rest of the site alone. A short trail, well laid, outperforms a wide one nobody follows.

Get the plumbing right underneath it. Google's hreflang documentation offers three methods, HTML tags, HTTP headers or sitemap entries, and asks you to choose one rather than mix them. The bidirectional rule is the one most sites break: if two pages do not both point to each other, the tags will be ignored. Add x-default for everyone your set does not match.

The multi-regional guidance repays a second read. Country-code domains give clear geotargeting but cost more and target one country each; subdirectories are easy to set up and low maintenance on the same host; URL parameters are explicitly not recommended. It also warns against translating only your boilerplate while the bulk of the content stays in one language, precisely the half-measure a rushed English folder becomes.

This is where the term generative engine optimization earns its keep, and where it stops. As a name for writing pages a retrieval system can quote, it is useful. As a licence to spin up language variants at scale, it is a spam policy violation with a fashionable label on it. The AI features documentation is blunt on the same point: there are no additional requirements to appear in AI Overviews or AI Mode, and no special optimisations necessary. Multilingual SEO does not get an exemption from that.

Section 05

What to build from a Maltese desk#

This is home ground for us. folkfox works out of Malta into a dozen European markets, where a client's Dutch page, French page and Maltese page all matter to real customers and none of them is the page a model reaches for first. Our whole multilingual SEO brief now fits in one breath: one genuinely written English version of your highest-intent pages, written for a machine's research step rather than translated for a human's browse.

Measurement is where most multilingual website seo programmes go quiet, because the old numbers do not move. Start with Microsoft, which has the cleanest view. The Bing Webmaster Tools AI Performance report, in public preview since February, shows when your site is cited in AI-generated answers across Microsoft Copilot, AI-generated summaries in Bing and select partner integrations, and which URLs are referenced over time.

Read it with its limits in view. Search Engine Land's write-up lists total citations, average cited pages, grounding queries and page-level citation activity, then notes these metrics only reflect citation frequency and say nothing about ranking, prominence or how a page contributed to an answer. Useful, not decisive.

Google gives you less. Its AI features documentation confirms AI feature impressions are folded into overall search traffic in Search Console under the Web search type, with no separate breakdown to isolate. So compare English URLs against local ones inside one property, on the same queries, and accept a shape rather than a number. Multilingual SEO reporting has to live with that for now.

One last piece of housekeeping. FAQ markup no longer earns a rich result and is absent from Google's current structured data gallery. Keep writing the questions, because self-contained answers are exactly what a fan-out lifts, but do not sell anyone a rich result that stopped existing.

If you want this built rather than described, that is folkfox SEO and GEO services, alongside content marketing for the writing itself. The pattern bites hardest in licensed sectors, where a machine paraphrasing your terms is a compliance problem rather than a traffic one, which is why we run it into iGaming marketing and FinTech marketing first.

Twelve months from now this is either settled science or a curiosity some model update quietly erased. Either way the cost of finding out is ten pages of honest English writing, a cheap bet on a clear signal. That is what good multilingual SEO looks like when the quarry has changed its habits and nobody has told the hunters yet.

Questions

Frequently asked questions#

Does ChatGPT really search in English when I ask in another language?

Often, yes. Peec AI analysed over 10 million ChatGPT searches and over 20 million query fan-outs, filtered so user location matched query language, and found 43% of research steps ran on the English-speaking web. Turkish prompts switched to English 94% of the time, Spanish 66%, and no non-English language fell below 60%. The answer still comes back in your language.

Should multilingual SEO mean translating my whole website into English?

No. Google's spam policies name generating many pages through automated translation, where little value is added, as scaled content abuse. The defensible move is a small number of genuinely written English pages covering your highest commercial intent, not a machine-translated mirror of everything you publish.

How much lift does an English folder actually give?

It varies by engine and the panels are small. Search Engine Land measured a 52% aggregate difference for Copilot across 272 Bing properties, a 28% aggregate lift for Google AI, and a 122% uplift for ChatGPT from only 26 properties. Treat the ChatGPT figure as indicative rather than proven.

Does hreflang still matter if assistants prefer English?

Yes, and arguably more. Hreflang is what stops your English page from cannibalising the local one in ordinary search. Google requires the annotations to point at each other bidirectionally or they are ignored, asks you to pick one implementation method, and supports x-default for languages your set does not cover.

A larger study across 400 questions and five countries shows that language changes which sources are cited.

Which crawler decides whether ChatGPT can see my English page?

OAI-SearchBot. OpenAI's documentation separates it from GPTBot: OAI-SearchBot surfaces websites in ChatGPT's search features, while GPTBot crawls content that may be used to train the foundation models. Blocking GPTBot for training reasons does not have to mean blocking search retrieval, and many sites block both by accident.

How do I measure whether the English page is being cited?

Use Bing Webmaster Tools' AI Performance report for Copilot citations, cited URLs and grounding queries. Google folds AI feature impressions into overall Search Console traffic under the Web search type, so multilingual SEO reporting has to compare English against local URLs on the same queries inside one property and read the shape rather than an exact number.

Keep reading

Read more on this topic#

Want the English page that earns the citation?

folkfox builds multilingual SEO around the ten pages a research step actually reaches for, wires the hreflang underneath them, and reports on citations rather than hope.