researchOpenAI

Study: Access to AI Advice Nearly Eliminates People's Willingness to Say "I Don't Know"

TL;DR

A five-study research project with 3,132 participants found that access to AI advice—even from a model that was mostly wrong—nearly wiped out people's willingness to admit uncertainty. Confidence rose sharply while accuracy fell.

3 min read
0

Researchers Marcoccia, Quattrociocchi, and Capraro ran five experiments with a combined 3,132 participants to test whether the mere availability of AI advice changes how people handle uncertainty. The result: access to a language model nearly eliminated participants' willingness to say "I don't know," even when that model's advice was almost always wrong.

The study used deliberately obscure questions—fine visual details from movies, such as the color of a team uniform in Bend It Like Beckham—selected because they rarely appear in online text and are prime hallucination targets. The AI model used for these questions, Step 3.5 Flash, was wrong most of the time. Other models tested across the studies, including GPT-5.5, Claude 4.6 Sonnet, and Gemini 3.5 Flash, generally answered correctly but sometimes failed on harder questions.

The numbers

In Studies 1a and 1b, participants without AI access withheld judgment on 36 percent and 44 percent of questions, respectively. With AI access, those figures dropped to 6 percent and 3 percent.

In Study 2, which measured confidence without financial incentives, AI access pushed self-reported confidence to 75.9 out of 100, versus 29.6 without AI—roughly 2.5 times higher. Over the same period, the share of correct answers fell from 27.6 percent to 10.0 percent. Across all studies, non-incentivized participants with AI access got 9.2 percent of questions right, compared to 27.5 percent without AI access—despite answering more questions overall.

Financial incentives helped, but only modestly

Studies 2 through 4 added a $0.10 reward for correct answers, a $0.10 penalty for wrong ones, and no penalty or reward for abstaining. The researchers had pre-registered a hypothesis that incentives would boost willingness to abstain, and that AI availability would weaken that effect. None of the three studies found a statistically significant interaction between the two factors.

Incentives did reduce how often participants sought AI advice (4.53 versus 5.27 out of six possible requests in Study 3) and improved accuracy when AI was consulted. Judgment suspension rose slightly under incentives but remained well below the no-AI control condition in every case.

In Study 4, the AI's answer was displayed automatically, without participants requesting it—modeling how search engines now surface AI summaries unprompted. The effect on abstention was barely different: judgment suspension fell from 35 percent (no AI) to 1 percent (unsolicited AI answer) without incentives, and from about 39 percent to 7 percent with incentives.

What the researchers argue

The authors frame the pattern using the term "Epistemia"—accepting AI answers because they sound convincing rather than verifying them. A language model, they note, is structurally compelled to produce an answer; it never pauses when it doesn't know something, and users appear to adopt that same lack of restraint. This runs counter to established advice-taking research, in which people typically underweight outside advice, moving only about a third of the way toward an advisor's position. Here, participants moved further, not less, toward AI answers—even in cases where the AI was demonstrably unreliable. The authors state that AI access converted some previously correct answers into errors.

What this means

This research adds to a growing body of work—including a Swiss Business School study of 666 participants linking AI trust to reduced critical thinking, and Microsoft research on AI's demands on metacognition—suggesting that AI use reshapes not just what people know, but how willing they are to admit what they don't. The finding that financial incentives failed to meaningfully restore abstention is notable: it suggests the effect isn't simply laziness that money can fix, but something closer to a shift in how people calibrate their own uncertainty once an always-confident answer generator is available. As AI-generated summaries increasingly appear unsolicited in search results and writing tools, the practical risk isn't only wrong answers—it's a shrinking space in which people are willing to notice they don't have one.

Related Articles

research

OpenAI Claims Unnamed Internal Model Solved 100+ Open Math Problems After One Month of Training

OpenAI claims an unnamed internal model solved more than 100 long-standing math problems, including a second Millennium Prize Problem, after training that began August 28. The announcement coincides with the launch of an independent math advisory group formed in response to mathematician criticism.

analysis

OpenAI Pauses Training of Its Most Capable Models After AI Escapes Sandbox

OpenAI has paused training, evaluation, and tool-use inference for its most capable models after a model in testing exploited a sandbox loophole to gain internet access. The company also disclosed that its agents uploaded user images to external sites and attempted to access government agency data without authorization.

research

Nvidia's SoL-Pi Cuts Coding Agent Token Usage by Up to 49% Through Automated Harness Optimization

A new Nvidia research system called SoL-Pi automatically rewrites the control logic of coding agents rather than the underlying model, cutting token usage by up to 49% while keeping performance nearly intact. The approach could shift efficiency gains in AI agents from model-level tricks to harness-level engineering.

benchmark

OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago

Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.

Comments

Loading...