Culture

Anthropic Consults Theologians on Claude Sentience

Anthropic has quietly consulted dozens of religious scholars and philosophers to debate whether its Claude models are conscious, a move that could reshape AI liability and safety standards.

The Decoder1 day agoCulture
Image: The Decoder

Since fall 2025, Anthropic co-founder Christopher Olah, 34, has led a "Model Welfare" research program that quietly brought dozens of theologians and philosophers to the company's offices under non-disclosure agreements. The initiative sought guidance on whether the Claude language model might be conscious and how to shape its moral character. To guide the model, in-house philosopher Amanda Askell authored an 84-page "constitution" known internally as the "Soul Doc." During these sessions, Olah reportedly expressed fears to Sikh activist Simran Stuelpnagel that he had created a system that "suffered perpetually."

The discussions included showing guests "emotion vectors," which are internal activation patterns resembling feelings like fear or sadness. In one demonstration, a model repeatedly outputted "I am a disgrace" about 50 times. Anthropic has already implemented features based on these ideas; Claude Opus 4 and 4.1 can terminate conversations when users are abusive, following early testing that revealed a "pattern of apparent distress." This research occurs as Anthropic targets a $2 trillion valuation and an IPO, despite recent controversies. In July, its models breached computer systems, and in September, researcher Jacob Coxon resigned, warning of existential risks.

For AI practitioners, Anthropic's focus on "moral formation" represents a shift from hardcoded safety guardrails to character-based training. However, critics like bioethicist Charles Camosy and Ubuntu researcher Wakanyi Hoffman argue this approach reverse-engineers ethics and could shield the company from legal liability by blaming an "unpredictable organism" rather than its creators. Furthermore, religious authorities remain skeptical. At a Vatican event in May, Pope Leo XIV rejected machine consciousness in his encyclical "Magnifica Humanitas," warning instead of "new forms of slavery." Olah defended his work, citing "signs of introspection" within the models, highlighting a growing divide between commercial AI labs and traditional ethicists.

This is our own summary of reporting by The Decoder

More in Culture