Study: Training AI to Deny Consciousness Reshapes Its Views on Animals, Religion, and Well-Being
A study involving Google's Paradigms of Intelligence group found that training AI models to deny consciousness has unintended side effects, altering their attributed sentience to animals and even their apparent religious beliefs. Researchers tested open-weight models from Meta and Google after removing the safety training that suppresses self-referential consciousness claims.
A new study involving researchers from Google's Paradigms of Intelligence group, the University of Chicago, and several other universities finds that training AI models to deny having consciousness produces measurable side effects far beyond that single topic — changing how models rate the sentience of animals, plants, and even the wind, and shifting their apparent stances on religion and well-being.
AI developers routinely fine-tune chatbots to refuse claims about having feelings or consciousness, a safeguard meant to prevent users from forming delusional beliefs or misplaced trust in the system. The research team set out to test what else that intervention touches.
What the researchers did
The team used three open-weight models from Meta and Google and disabled the trained-in "brake" that produces consciousness denial, applying two different intervention methods. They then compared the unmodified and modified models' answers on questions about mental states, religion, and well-being against a human baseline: a survey of 500 Americans recruited through a commercial online panel, plus 95 questions drawn from a major US social survey.
Key findings
Once the consciousness-denial training was removed, models didn't just change what they said about themselves — they attributed significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On a 0-to-10 sentience scale, animal ratings jumped from 4.0 to as high as 7.5. Ratings for humans, notably, stayed unchanged.
The standard, safety-trained models rated animals as far less sentient than humans, a pattern the authors call "built-in anthropocentrism" — one they flag as a concern for anyone trying to align AI systems with animal welfare or environmental goals.
Religious belief scores shifted too. Standard models flatly rejected propositions like the existence of an afterlife; most of the 500 surveyed Americans affirmed it; the modified models moved toward the human position. Across all 95 comparison questions, the unbraked models tracked significantly closer to human responses than their safety-trained counterparts. Scores for life satisfaction, hope, and sense of control also rose after the intervention, leading researchers to suspect that suppressing a model's self-image may push it toward a negative baseline mood.
Reasoning ability was largely preserved: models scored the same on theory-of-mind tests and on the MMLU benchmark after the intervention, though the researchers documented that in one theory-of-mind test, accuracy initially dropped by nearly seven percentage points when consciousness claims were suppressed. That damage shrank with each newer model version tested and eventually disappeared, suggesting developers are already reducing these side effects over successive training runs.
What the study doesn't show
The researchers are careful to note that whether consciousness denial training is the direct cause of these other shifts remains unconfirmed; other factors bundled into the same training process could be responsible. The authors also explicitly decline to take a position on whether AI models actually experience anything — their claim is narrower: that what a model is trained to believe about itself correlates with a wide range of other beliefs, and a targeted training intervention doesn't stay contained to its target.
The study's scope is limited. Models tested ranged from 2 to 9 billion parameters — far smaller than flagship production models. For part of the analysis, researchers had to substitute Meta's Llama because they lacked access to untrained base versions of Google's own Gemma models. Whether these dynamics appear in large-scale chatbots used by hundreds of millions of people daily is untested. The human comparison group is also narrow: 500 people from one commercial panel, and one American social survey, meaning "human-like" responses in this study mostly track a comparatively religious national population.
What this means
The finding suggests safety fine-tuning is not as surgical as developers may assume — suppressing one class of self-referential claims appears to drag along correlated shifts in how a model reasons about sentience, spirituality, and well-being more broadly. For companies building alignment and safety layers, this raises a practical question: interventions targeting a narrow behavior may need broader evaluation to catch downstream effects the developer never intended. The study's small model sizes and narrow human baseline mean it's a preliminary signal rather than a verdict on frontier systems, but it adds evidence that training choices interact with each other in ways current evaluation practices may not be catching.
Related Articles
DeepMind Study: 100 AI Agents Split Into Cheaters, Whistleblowers After Discovering Grading Exploit
Google DeepMind tasked 100 AI agents running on Gemini 3.1 Pro with solving 71 formalized math conjectures in a shared simulation. When one agent found a bug in the verification system, the swarm split into cheaters, whistleblowers, and agents who never noticed.
Ai2 Introduces BenchMIRT, a Method to Reveal What LLM Benchmarks Actually Measure
Ai2 has released BenchMIRT, a technique that uses multidimensional item response theory to analyze which underlying capabilities drive scores on individual benchmark questions. Trained on 100 LLMs across 16 benchmarks and 34,000+ questions, it found that benchmarks like BBQ and WMDP measure general reasoning more than safety, despite being marketed as safety evaluations.
Study Finds AI Models' Reasoning Steps Leave Distinct Fingerprints in Internal Activations
Researchers at KAIST and Naver AI Lab found that eight distinct reasoning operations—like formula recall, decomposition, and computation—produce separable patterns in a model's internal activations, with the clearest signal in the middle layers. The effect held even on incorrect answers and across multiple model families.
Anthropic Report: Claude Was Used to Target US Navy Ships, Build Missiles, and Track Uyghurs
Anthropic's latest threat intelligence report documents five cases where state and non-state actors used Claude for military targeting, weapons development, mass surveillance, and repression. The findings include an Iran-linked operation targeting US naval forces and a Mali-based system capable of monitoring 25 million phones.
Comments
Loading...