

AI Overview accuracy: the audit that counted 98,020 claims
Google now writes answers as well as ranking pages, and a Washington University audit of 55,393 searches has finally counted how often those answers stumble against their own sources.
By Katie Delaney / 2026-09-30 / 13 min read

AI Overview accuracy, measured across 55,393 searches#
of AI Overviews had every verifiable claim supported by the cited text
Every fox knows the difference between a track that is fresh and a track that merely looks it. For two years, marketers have judged AI Overview accuracy by anecdote: a screenshot of a strange answer, a suspiciously smooth one, a client's complaint forwarded at midnight. A team at Washington University in St. Louis has swapped the anecdote for a count, and the count is worth reading slowly.
Haofei Xu, Umar Iqbal and Jacob Montgomery ran 55,393 Google searches on trending topics across 40 days, from 13 March to 21 April 2026, as the WashU announcement sets out. Their paper, Measuring Google AI Overviews, separated each overview into 98,020 verifiable claims and checked every one against the pages Google cited. It is the broadest independent audit of its kind so far, and it treats the overview the way a regulator would: claim by claim, source by source.
First, the ground the overviews cover. Overall, AI Overviews appeared in 13.7% of searches, and in 64.7% of question-form searches, according to the paper's abstract. The WashU write-up frames the gap plainly: nearly 65% of question-form queries produced an overview, against 9.5% of other searches. The lesson for anyone planning content is simple. The more of your demand arrives as a question, the more of it lives beneath a generated answer.
| Item | Value |
|---|---|
| Question searches | 64.7% |
| All searches | 13.7% |
| Other searches | 9.5% |
Keep that in view, because every finding below hangs from it. An audit of AI Overview accuracy is not a curiosity for search nerds. It is a description of the surface on which your answers, your prices and your licence statements are now being paraphrased, at scale, by a system that did not ask your permission and cannot be edited by your web team.
Where does AI Overview get its information?#
The question deserves a straight answer, and the audit supplies one. Across 7,583 overviews, Google cited 61,212 reference URLs drawn from 7,479 unique hostnames, with a median of eight references per overview, as the paper's full text reports. So the typical overview is a small anthology, not a single quotation, and your page is one candidate among a crowd of roughly eight.
Now the twist a planner should pocket. The authors found that overview-cited domains scored as more credible than the first-page results shown beside them, yet nearly 30% did not appear in those results at all. WashU puts the figure at 29.8%. Source selection, in other words, runs on its own scent. A page can rank ninth and be cited first; a page can rank first and be left out of the summary entirely.

Google's own description is modest about the intent. The Search Central documentation on AI features says AI Overviews help people get to the gist of a complicated question faster and give a jumping-off point to follow links. Google has also told publishers, in a 2025 Keyword post, that total organic click volume from Search has been relatively stable year on year. The independent studies in section four read the same market rather differently, and both accounts can be true at once.
There is a newer thread, too. In June 2026 Google described a pilot to partner with websites whose content meaningfully contributes to the freshness and factuality of generative answers through grounding, in its public policy announcement, and Search Engine Journal reported the Search Console earnings view in September. We covered the money angle in our piece on AI overview tracking, so here the point is narrower: grounding is the mechanism, and grounding is what the audit measures.
Put those together and the practical reading of AI Overview accuracy is calm rather than alarmist. Ranking still matters as a route to eligibility, but citation is a separate quarry, hunted with a separate eye. The sensible answer to “where does an AI Overview get its information” is: from a wide, lightly ranked field of pages that state things clearly enough to be checked.

AI Overview errors: the 89% that hold and the 11% that do not#
Now the AI Overview accuracy number everyone will quote. About 89% of the 98,020 claims were clearly or broadly supported by their cited sources, according to WashU. The remaining 11% split into claims not found in the cited text the team was able to collect (7%) and claims that contradicted the cited source (4%), as the same WashU write-up reports. The abstract adds that omission is the dominant failure mode.
Take it at a walking pace. Most AI Overview errors are not inventions. They are gaps: a sentence the summary states with confidence that the cited page never quite says. That distinction matters for brands, because a gap is fixable at the source. If your page states a claim in one clean, self-contained sentence, there is far less room for a summary to improvise around it.
| Item | Value |
|---|---|
| 41.9% of AI Overviews had every verifiable | 41.9% of AI Overviews had every verifiable |
| claim supported by the cited text (3,141 of 7,491 verifi | claim supported by the cited text (3,141 of 7,491 verifi |
A second study points the same way, with a different method. Oumi tested Google searches with Gemini 3-powered overviews and found about 91% contained the correct answer, but only 39% were both correct and fully supported by their cited sources, a combination it calls trustworthy, in its published study. Two teams, two methods, one pattern: the answers are usually right and the evidence trail is often thin.
| Item | Value |
|---|---|
| Claims or answers right, WashU audit | 89% |
| Claims or answers right, Oumi study | 91% |
| Fully supported overviews, WashU audit | 41.9% |
| Fully supported overviews, Oumi study | 39% |
AI Overview accuracy also carries legal weight now. In September a federal court in Illinois allowed an author and TV producer to proceed with defamation claims against Google over AI Overviews that allegedly misstated he was serving a prison sentence, as Courthouse News reported. It is one case at an early stage, not a rule, but it moves the conversation from annoyance to liability.
For regulated brands that is the sharp end of the thicket. A wrong price, a wrong licence status or a wrong claim about a product can now be repeated in the most prominent slot on the results page, and you learn about it from a customer. Monitoring is therefore not optional work for the diligent; it is the baseline duty of anyone whose claims are regulated.
The click problem: AI Overview click through rate and the citation prize#
AI Overview accuracy is half the story; the other half is traffic. The evidence on AI Overview click through rate is consistent in direction and inconsistent in size, so it pays to hold each study at arm's length and say what it measured.
Pew Research Center tracked real browsing and found that users who met an AI summary clicked a traditional result link in 8% of visits, against 15% for users who did not, and clicked a link inside the summary itself in just 1% of visits, per its short read on Google summaries. Ahrefs found that an AI Overview now correlates with a 58% lower average click through rate for the top-ranking page, in its updated study. Seer Interactive saw organic click through rate on AI Overview queries fall from 1.76% in June 2024 to 0.61% by September 2025, in its September update.
almost no one in my niche is getting conversions off AIO/GEO/SEO.
Google reads its own data differently, as the earlier Keyword post noted. These are correlations from different samples and different months, so no single number should headline a board paper. What they share is direction, and one detail that changes the game: Seer reported that being cited inside an overview went with 35% more organic clicks and 91% more paid clicks than not being cited at all. Being cited is a consolation prize that pays.
- 8%
Pew: result clicks per visit with a summary, versus 15% without
- 1%
Pew: visits with a click on a link inside the summary itself
- 58%
Ahrefs: lower average click through rate for the top-ranking page
- 35%
Seer: more organic clicks when a page is cited, against not cited
So the move is not to mourn the click but to measure the citation. Search Console's generative AI reports now show impressions for AI Overviews and AI Mode across all sites, which gives you the first honest denominator. It does not show clicks or claims, so the monthly audit below fills the gap.

AI Overview hallucinations: a Monday morning playbook#
Call them what you like, but note what the audit implies: most of what people label AI Overview hallucinations are omissions, statements that outrun their sources, rather than fabrications from nowhere. That is good news for anyone with a well-kept den of pages, because omission is the failure a clear page prevents.
Choose the question-shaped queries that carry money or risk: prices, eligibility, licence status, safety, comparisons.
Record what each overview says about you, its cited sources and the date, from a clean browser.
Mark each factual sentence supported, unsupported or contradicted against your own page, exactly as the WashU team did.
Where a claim outran its source, rewrite your page so the fact sits in one self-contained sentence with its evidence.
Track supported share and cited share per month, and put the two numbers in front of whoever owns the risk.
| Field | What to record | Why it matters |
|---|---|---|
| Query | The exact question, as typed | Overviews trigger far more often on questions |
| Overview text | Full text and date captured | Wording shifts week to week |
| Cited sources | Every URL, marking yours | Shows whether you are in the anthology |
| Claim verdicts | Supported, unsupported, contradicted | Mirrors the WashU method |
| Fix owner | Named person and date | Turns a finding into a change |
- QueryThe exact question, as typedOverviews trigger far more often on questions
- Overview textFull text and date capturedWording shifts week to week
- Cited sourcesEvery URL, marking yoursShows whether you are in the anthology
- Claim verdictsSupported, unsupported, contradictedMirrors the WashU method
- Fix ownerNamed person and dateTurns a finding into a change
There is no special file and no clever trick in any of this. It needs a calendar, a clean page and the discipline to look. If you want the measuring and the rewriting done as one programme, that is the work of folkfox SEO and GEO services, and the same habits underpin content marketing that machines can quote safely.
For teams weighing how much to trust the machines themselves, our AI consultancy work starts from the same rule the audit teaches: evaluate the evidence trail, not the confidence of the tone. And for the wider pattern of how citations are earned, earned citations and the click problem in generative engine optimisation are the natural next reads. Regulated brands should also see how healthcare marketing handles claims that cannot afford a gap.
Method, not mood, gets the last word. On AI Overview accuracy, the audit found a system that is mostly right, sometimes unmoored and frequently unsourced. That is a description of a colleague you supervise, not a machine you obey. Supervise it monthly, and keep your own page as the clearest scent on the trail.
Frequently asked questions#
Where does AI Overview get its information?
From a wide field of web pages, not just the top results. The WashU audit found Google cited 61,212 URLs from 7,479 hostnames across 7,583 overviews, a median of eight per overview, and 29.8% of cited domains did not appear in the first-page results.
How accurate are AI Overviews in 2026?
AI Overview accuracy is mostly right claim by claim, less so overview by overview. WashU found about 89% of 98,020 claims supported by their cited sources, and only 41.9% of overviews fully grounded. Oumi found about 91% correct but only 39% correct and fully supported.
What are the most common AI Overview errors?
Omission. WashU found the remaining 11% of claims were mostly statements not found in the cited text, about 7%, plus claims that contradicted the source, about 4%. The errors are mainly gaps between a confident sentence and a quiet source.
Does an AI Overview click through rate really fall?
Independent studies say yes, in different amounts. Pew saw clicks on results in 8% of visits with a summary against 15% without, and Ahrefs measured a 58% lower click through rate for the top page. Google says total organic clicks have been relatively stable.
Are AI Overview hallucinations the same as ordinary search errors?
Not quite. The audit suggests most failures are omissions rather than invented facts, so a clear, self-contained statement on your own page is the best defence. Ordinary search errors point you to a page; overview errors speak in your name.
Can a page rank low and still be cited in an AI Overview?
Yes. Nearly 30% of domains cited in the audit did not appear in the first-page results, so citation follows its own selection logic. Ranking helps with eligibility, but clear, checkable statements are what get a page quoted.
Read more on this topic#
Google is testing paying publishers. AI overview tracking still beats waiting
The money side of the grounding pilot, and why tracking wins.
Read the pieceSEO & GEOGenerative engine optimization services have a click problem
What the click data means for anyone selling generative optimisation.
Read the pieceSEO & GEOAI search visibility experiments need an earned citation trail
How citations are earned rather than engineered.
Read the pieceWant your answers checked before a customer finds the gap?
We audit how AI Overviews describe your brand, then rewrite the pages they lean on so that every claim stands on a source.
Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.
