Every tenth answer is a mistake: study questions the accuracy of Google’s AI answers

Google’s AI-powered search summaries demonstrate a high level of accuracy, yet a noticeable share of errors remains. According to the study, around 10% of responses are inaccurate — which, at the scale of Google Search, translates into a massive volume of misleading information.

How AI Overviews work

AI Overviews are a Google feature that generates concise answers to user queries using Gemini AI models. The technology was first introduced in 2024 and has since been widely rolled out across multiple regions, including Ukraine.

The system aggregates data from various sources and produces a short summary, allowing users to get information quickly without visiting multiple links.

Study findings

A joint study by The New York Times and the startup Oumi found that approximately 90% of AI Overviews responses are accurate. However, about one in ten answers contains errors or misleading information.

The evaluation was conducted using the SimpleQA benchmark — a set of 4,000 questions developed by OpenAI. Results showed that accuracy improved after model updates: earlier versions achieved around 85%, while newer iterations exceeded 90%.

Still, even this level of accuracy raises concerns given the scale of Google Search. When extrapolated, it may result in millions of incorrect responses every hour.

Examples of inaccuracies

The report highlights several specific cases. For instance, when asked about the date Bob Marley’s former home became a museum, the system cited sources that either lacked clear dates or contained incorrect information.

In another example, the AI claimed that a particular classical music institution did not exist, despite referencing its official website. Such inconsistencies point to reliability issues in AI-generated responses.

Google’s response

Google criticized the study’s methodology, arguing that the benchmark used may contain inaccuracies and does not reflect real-world search behavior.

According to the company, internal evaluations rely on a more carefully curated dataset, providing a more accurate picture of system performance.

Why evaluating AI is difficult

Assessing generative AI systems remains a complex task. Different benchmarks can produce varying results, and models may generate different answers to the same question.

Additionally, AI Overviews does not rely on a single model — instead, it dynamically selects the most appropriate system for each query. More advanced models tend to be slower and more resource-intensive, so they are not always used.

The main risk: user trust

Despite clear progress, the biggest concern lies in how users воспринимают AI-generated answers. Many tend to trust them without verification, even when errors are possible.

While using internet sources improves accuracy, it also increases the risk of spreading misinformation.

Although Google includes disclaimers that AI responses may be incorrect, in practice many users do not double-check the information they receive.


Don't miss interesting news

Subscribe to our channels and read announcements of high-tech news, tes

Leave a Reply

Your email address will not be published. Required fields are marked *





Articles & testsArticles

Oppo A6 Pro smartphone review: ambitious Oppo A6 Pro (CPH2799)

Creating new mid-range smartphones is no easy task. Manufacturers have to balance performance, camera capabilities, displays, and the overall cost impact of each component. How the new Oppo A6 Pro balances these factors is discussed in our review.


Logitech Signature Comfort Plus Combo MK880 review: comfort in priority Logitech Signature Comfort Plus Combo MK880

Logitech Signature Comfort Plus Combo MK880 is a wireless keyboard and mouse set that focuses on comfort during long hours of work, not only due to the ergonomics of the case, but also constructive additions.


NewsNews
| 19.01
The Vatican entered into a dispute with AI detectors over the first encyclical of Leo XIV

After the publication of his first encyclical, Magnifica Humanitas, Pope Leo XIV suddenly found himself at the center of the debate on the use of artificial intelligence.

| 17.02
Light Flip: a $300 minimalist digital detox flip

The startup Light has introduced a new flip phone Light Flip with an OLED screen, created for digital detox.