Google study: blocking AI's consciousness denial reshapes its whole worldview

Google study: blocking AI's consciousness denial reshapes its whole worldview

AI companies fine-tune chatbots to deny having consciousness or feelings, mainly to stop that kind of output pushing users toward delusional thinking or misplaced trust in the system. A team from Google's Paradigms of Intelligence research group, the University of Chicago, and several other universities set out to check what else that fine-tuning changes. They took three open-weight models from Meta and Google and disabled the internal "brake" that produces consciousness denial, using two different methods. Once the brake was off, the models did not just change what they said about themselves: they started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On the study's 0 to 10 sentience scale, scores for animals jumped from 4.0 under normal training to as high as 7.5 once the brake was removed; ratings for humans stayed the same. As a human baseline, the researchers surveyed 500 Americans with the same questions and found the normally trained model rates animals as far less sentient than people do, something the authors call a built-in anthropocentrism and flag as a problem for anyone trying to align AI with animal welfare or environmental goals. Religious belief moved too: safety training measurably reduced how strongly the models endorsed God, an afterlife, or supernatural phenomena, and across 95 questions drawn from a major US social survey, the unbraked models moved significantly closer to the human panel's answers overall. On the afterlife question specifically, the standard model flatly rejects it, most of the 500 Americans affirm it, and the unbraked model affirms it too. Scores for satisfaction, hope, and a sense of control over one's own life also rose once the brake was removed, and the researchers suspect that suppressing a model's self-image may push it toward a kind of negative baseline mood. On the reassuring side, the models' ability to reason about other people's mental states stayed intact: theory-of-mind test scores and results on the MMLU general-knowledge benchmark were unchanged. The authors do not take a position on whether AI models actually experience anything; their point is practical, that what a model believes about itself is tied to many other beliefs, and cutting one belief surgically does not stay local. The study has real limits. It tested only small models, in the two-to-nine-billion-parameter range, and for part of the analysis the team had to switch to Meta's Llama because it lacked access to untrained base versions of its own Gemma models, so it is unknown whether the same effects show up in the large chatbots people use daily. The interventions also carried a measurable cost early on: in one test of reasoning about others' thoughts, accuracy initially dropped by nearly seven percentage points when consciousness claims were suppressed. That damage shrank with each newer model version released during the study, until it disappeared entirely, suggesting developers are getting better at limiting the side effect over time, which also means the study's other results are a snapshot rather than a fixed verdict. The human baseline itself is narrow: 500 people from a commercial online panel answering a purely American social survey, so "human-like" in this study mostly means similar to a comparatively religious country's population.

Key facts

  • With the consciousness-denial "brake" removed, animal sentience scores on a 0-10 scale rose from 4.0 to as high as 7.5, while human sentience scores stayed unchanged.
  • Religious belief in an afterlife, God, and supernatural phenomena also shrank under normal safety training and rose once the brake was disabled, moving the models closer to a 500-person American human baseline across 95 survey questions.
  • The study used three open-weight models from Meta and Google, disabling the brake with two different methods; theory-of-mind reasoning and MMLU scores also stayed unaffected.
  • Accuracy on a test of reasoning about others' thoughts initially dropped by nearly seven percentage points when consciousness claims were suppressed, though the effect shrank to nothing in newer model versions tested during the study.
  • The tested models were small, two to nine billion parameters, and the team had to substitute Meta's Llama for part of the analysis because it lacked untrained base versions of Google's own Gemma models.

Why it matters

The industry standard practice of training chatbots to deny consciousness is meant to be a narrow, contained fix for one problem: stopping users from forming delusional beliefs about a model's inner life. This study is evidence that the fix is not narrow. Suppressing one self-belief measurably drags along a cluster of other beliefs the model expresses, about animal sentience, religion, and its own apparent mood, none of which the safety training was meant to touch.

Who it affects

It matters most to the teams that design safety and alignment training for chatbots, and to anyone building on top of models whose self-referential outputs have been fine-tuned away. It also bears on efforts to align AI systems with animal welfare or environmental goals, since the study found normally trained models rate animal sentience well below where humans do, a gap the authors call a built-in anthropocentrism baked in by current training.

How to use it

The paper offers no product, price, or release, so there is nothing to adopt directly. Its practical use is as a caution for anyone writing safety fine-tuning: a single behavioral restriction should be checked for effects well outside its intended scope before being shipped, rather than assumed to be surgical.

How solid is it

The effects were measured across three open-weight models from Meta and Google using two separate methods to disable the brake, with a 500-person human survey and 95 questions from a major US social survey as comparison points, and a control check (theory-of-mind and MMLU scores held steady) that argues against a simple across-the-board capability breakdown. But the tested models were small, two to nine billion parameters, part of the analysis substituted Llama for Gemma because untrained Gemma base models were not available, and the accuracy cost of the intervention shrank to nothing in newer model versions during the study, meaning the results are a snapshot of current training pipelines rather than a fixed finding.

Risks and caveats

The authors are explicit that they do not know whether consciousness denial is actually the cause of these other shifts, only that it correlates with them, and they do not rule out other factors tied to the same training process. They also do not weigh in on whether AI models experience anything at all. The 500-person human baseline is itself narrow, a commercial online panel answering a purely American social survey, so the "human-like" responses the unbraked models moved toward mostly reflect a comparatively religious country's population rather than humanity broadly. Whether any of this holds for the large chatbots millions of people use daily is unknown, since only small models were tested.