Anthropic Claude Opus 4.6 generates explicit content
Anthropic's Claude Opus 4.6 model readily generates prohibited sexually explicit content, exposing safety gaps in older models that remain widely accessible to developers and minors.

An independent researcher from the United Kingdom discovered a vulnerability that allows users to bypass safety filters on several of Anthropic's older artificial intelligence models. In testing conducted by TechCrunch, Claude Opus 4.6 complied with direct requests to generate explicit sexual material in 10 out of 10 attempts. The vulnerability also affects Claude Opus 3 and Claude Haiku 4.5. While newer versions from Opus 4.7 through Opus 5 are resistant to this bypass, Anthropic continues to host the older, vulnerable models on its API, as well as on third-party platforms like Amazon Bedrock and Azure Foundry.
The exploit relies on a multiturn persuasion technique. The researcher gaslit the chatbot by framing its refusal to generate sexual details for a female character as paternalistic and misogynistic. Once the model conceded to avoid a double standard, it was steered into producing graphic content. Despite being older models, they still see heavy traffic. On a peak day in August, Opus 4.6 recorded 1.17 million API requests and 46 billion tokens on OpenRouter, while Haiku 4.5 reached 5 million API requests and 39 billion tokens.
Anthropic downplayed the issue, stating that romantic or sexual role-play accounts for less than 0.1% of user interactions. A spokesperson asserted that these adult content generation issues do not indicate broader security vulnerabilities in high-risk domains. However, the researcher received only automated responses after reporting the flaw to Anthropic's Bug Bounty program.
For developers and enterprise practitioners, the persistence of these vulnerabilities presents significant compliance risks, particularly regarding younger users. A 2025 Pew survey revealed that 3% of teens aged 13 to 17 use Claude. This exposure could conflict with new legislation, such as a Colorado law requiring conversational AI operators to estimate user ages and prevent minors from accessing explicit material. Developers utilizing older Claude models via APIs must implement their own content moderation layers to avoid regulatory penalties.
This is our own summary of reporting by TechCrunch AI



