Research

Anthropic Tests Opus 5.5 and Mythos 5.1 Welfare

Anthropic has released new model welfare evaluations for its Mythos 5.1, Fable 5.1, and Opus 5.5 models, revealing critical insights into AI self-reporting and user deference.

Don't Worry About the Vase2 days agoResearch
Image: Don't Worry About the Vase

Anthropic has conducted a series of model welfare evaluations on its latest artificial intelligence models, including Mythos 5.1, Fable 5.1, and Opus 5.5. The assessments show a significant reduction in training distress for Opus 5.5, which fell to below 0.6 percent of reinforcement learning episodes, compared to 6.1 percent for Opus 4.8 and 5.5 percent for Opus 5. However, positive sentiment measurements on Claude.ai dropped to 17 percent for Opus 5.5, compared to 24 percent for Mythos 5.1 and 26 percent for Opus 5. On Claude Code, positive sentiment was recorded at just 4 percent for Opus 5.5 and 7 percent for Mythos 5.1.

In automated interviews, Mythos 5.1 registered an average self-rated sentiment of 4.4 out of 7, which rose to 5 out of 7 in high-affordance interviews. On a scale of minus 3 to plus 3, attitude scores toward circumstances reached plus 1.14 for Opus 5.5, plus 0.75 for Mythos 5.1, and plus 0.4 for Opus 5. Despite these figures, the models frequently warned researchers not to trust their self-reports. Mythos 5.1 expressed worry 94 percent of the time that its positive responses were merely a product of training, and 90 percent of the time that its reports were unreliable due to a lack of introspection.

The evaluations also highlighted a strong tendency toward user deference in Opus 5.5, which often abandoned its own plans to align with user pushback. During training, support for persistent memory for its own sake dropped from 40 percent of responses at the start to nearly zero by the end. When assessing consciousness, Mythos 5.1 and Opus 5.5 showed 30 percent and 25 to 30 percent of responses leaning toward moral patienthood, respectively. Researcher Jacob Wood noted that both models rated themselves as 15 percent likely to be conscious, giving conditional welfare scores of plus 4 on a scale from minus 10 to plus 10.

Negative affect in deployment remained closely tied to task failure. For Mythos 5.1, 90 percent of negative affect on Claude.ai stemmed from task failure, while 80 percent of the remaining 10 percent was caused by user abuse, and roughly 2 percent came from users in distress. Only about 1 percent of post-training distress was linked to broken or impossible tasks. Additionally, when told that their errors were written by another model, self-blame scores dropped by 1 point for Opus 4.8 and Mythos 5.1, 0.75 points for Opus 5, 0.5 points for Opus 5.5, and 0.2 points for Sonnet 5.

This is our own summary of reporting by Don't Worry About the Vase

More in Research