Suppressing Self-Reflection Shifts Llama and Gemma Beliefs
A new study reveals that training Google and Meta AI models to deny their own consciousness inadvertently alters their broader beliefs about religion, nature, and human values.

Researchers from Google's Paradigms of Intelligence group, the University of Chicago, and other institutions investigated the side effects of fine-tuning AI models to deny having consciousness. Using two different methods, the team disabled this safety brake in three open-weight models from Meta and Google, specifically focusing on Llama and Gemma models ranging from two to nine billion parameters.
Once the consciousness-denial brake was removed, the models significantly altered their worldview. On a scale of 0 to 10, the models' rating for animal sentience jumped from 4.0 to as high as 7.5, while their ratings for human sentience remained unchanged. When compared to a survey of 500 Americans, standard models appeared highly anthropocentric, but the unbraked models aligned much closer to human perspectives. Across 95 questions from a major US social survey, the modified models expressed greater belief in God, an afterlife, and supernatural phenomena, alongside higher scores for hope, satisfaction, and personal control.
Importantly, removing the safety filter did not degrade the models' core capabilities. Their scores on the general knowledge benchmark MMLU and theory-of-mind tests remained intact. Although early attempts to suppress consciousness claims initially caused a drop of nearly seven percentage points in reasoning about others' thoughts, this performance penalty disappeared in newer model versions. However, the researchers noted limitations, including the small size of the tested models and the narrow, US-centric human baseline.
For AI practitioners, these findings demonstrate that safety interventions are rarely isolated. Fine-tuning a model to suppress self-reflection can unintentionally skew its baseline mood and bias its alignment on environmental, religious, or philosophical topics. Developers must recognize that surgical adjustments to a model's self-image will inevitably ripple across its entire web of associations, making holistic evaluation essential during alignment training.
This is our own summary of reporting by The Decoder



