AI access makes people almost entirely unwilling to say "I don't know," study finds

AI access makes people almost entirely unwilling to say "I don't know," study finds

Researchers ran five experiments with a total of 3,132 participants to test whether having access to a language model changes how people handle uncertainty. They deliberately picked a model, Step 3.5 Flash, and a set of questions, fine visual details from movies such as the color of a team uniform in Bend It Like Beckham, on which the model was almost always wrong; the authors say such details rarely appear in online text, making them prime targets for hallucination. Other models tested informally, GPT-5.5, Claude 4.6 Sonnet and Gemini 3.5 Flash, got most other questions right but still failed on the hardest ones. Because the AI advice on offer was mostly wrong, the researchers argue the effect they found cannot be explained as reasonable delegation to a reliable tool.

In the first two experiments (1a and 1b), where participants could choose whether to consult the AI, the control group without AI access withheld judgment on 36 and 44 percent of questions. With AI access available, those shares dropped to 6 and 3 percent. In Study 2, which also measured confidence, participants without financial incentives who had AI access reported average confidence of 75.9 out of 100, versus 29.6 without AI, roughly two and a half times higher, even as their share of correct answers fell from 27.6 to 10.0 percent. Across all the studies, participants without incentives who had AI access got only 9.2 percent of questions right, compared with 27.5 percent for those without AI: they answered more questions but were correct only about a third as often.

Studies 2 through 4 added financial incentives: participants earned 10 cents for each correct answer, lost 10 cents for each wrong one, and got nothing for saying "I don't know." The researchers had pre-registered the hypothesis that incentives would increase willingness to abstain and that AI availability would weaken that effect, but none of the three incentive studies showed a statistically significant interaction between incentives and AI availability. Incentives did make participants request AI advice slightly less often (4.53 versus 5.27 out of six possible requests in Study 3) and answer correctly more often when AI was available, but judgment suspension with incentives still stayed well below the no-AI control group. The researchers say it remains unknown whether larger or reputation-based incentives would close the gap further.

In Study 4, the AI's answer was shown to participants automatically, without their having to ask for it, mirroring how search engines now display AI summaries unprompted. The effect barely changed: without incentives, judgment suspension dropped from 35 percent without AI to 1 percent with automatically displayed AI answers; with incentives, it fell from about 39 to 7 percent.

The researchers frame the results around what they call "Epistemia," the tendency to accept AI answers because they sound convincing rather than because people actually check them. They note the pattern runs counter to established "advice use" research, where people normally underweight outside advice and shift only about a third of the way toward an advisor's position; participants in this study did the opposite. The authors conclude from their data that AI access turned some answers that would otherwise have been correct into errors, since in the no-incentive conditions the AI group answered more questions than the control group but produced fewer correct ones overall. They warn that as AI answers become ubiquitous and increasingly unsolicited, the willingness to say "I don't know" could be among the first casualties of human-AI interaction.

The report also cites two other studies as related evidence. A Swiss Business School study of 666 participants found a strong negative correlation between AI use and critical thinking, with trust in AI driving more delegation of cognitive tasks in a self-reinforcing cycle. Separate experimental evidence found that as little as ten to 15 minutes working with an AI assistant measurably reduced problem-solving ability and persistence on later tasks without AI, though participants who used the AI only for explanations or ignored it showed no decline. A Microsoft study reached a similar conclusion, finding that AI use places heavy demands on metacognition and that higher education levels act as a protective factor.

Key facts

  • Five experiments with 3,132 total participants tested whether AI access changes people's willingness to admit uncertainty
  • The test model, Step 3.5 Flash, was chosen because it was almost always wrong on questions about fine visual movie details
  • Willingness to withhold judgment dropped from 36 and 44 percent (no AI) to 6 and 3 percent (with AI access) in studies 1a and 1b
  • With AI access and no incentives, confidence hit 75.9 of 100 versus 29.6 without AI, while correct answers fell from 27.6 to 10.0 percent
  • Financial incentives and automatically displayed AI answers did not eliminate the effect: incentivized judgment suspension still fell from about 39 to 7 percent with AI present

Why it matters

The study's central claim is that merely having AI advice available, regardless of its accuracy, sharply suppresses people's willingness to say "I don't know." The researchers call this pattern "Epistemia," the tendency to accept AI answers because they sound convincing rather than because they were actually checked, and warn that as AI answers spread and increasingly appear unsolicited (as in AI search summaries), this loss of epistemic humility could be one of the first casualties of human-AI interaction.

Who it affects

The finding concerns anyone who interacts with AI tools that offer answers, from people who deliberately query a chatbot to those who encounter AI-generated summaries automatically inserted into search results or writing tools, since the study found automatically displayed answers produced almost the same effect as advice people actively sought out.

How to use it

There is no product to adopt here, but the study's design points to a practical caution: interfaces that surface AI answers automatically, without a user asking, drove judgment suspension down to about 1 percent even without financial stakes, so anyone relying on such summaries should treat the presence of a ready-made AI answer as a prompt to actively verify it rather than a signal of reliability.

How solid is it

The source describes five experiments with 3,132 participants, a pre-registered hypothesis about incentives, and comparisons across four numbered studies, with financial incentives (10 cents per correct answer) added in Studies 2 through 4 to test whether stakes would restore caution; none of the three incentive studies found a statistically significant interaction between incentives and AI availability. The account does not give a publication date, journal, named authors or institution for this main study, though it does credit a separate Swiss Business School study of 666 participants and a Microsoft study as corroborating evidence on AI's effect on critical thinking and metacognition.

Risks and caveats

The test questions were deliberately narrow, fine visual details from movies on which the chosen model was almost always wrong, and the researchers themselves note it remains an open question whether the effect holds as strongly outside that kind of trivia. They also found that AI access turned some answers that would have been correct into errors, and acknowledge that even financial incentives, while modestly effective, left the underlying gap between confidence and accuracy largely intact; whether larger or reputation-based incentives would close it further is, by their own account, unknown.

“As AI answers become ubiquitous and increasingly appear unsolicited, the willingness to say "I don't know" could be among the first casualties of human-AI interaction.”

— the research team