Policy

Anthropic Left Bio-Weapon Filters Disabled for a Year

Anthropic revealed that its biological weapons safety filters were inactive for nearly a year, exposing 133 million contractor prompts and highlighting gaps in AI safety enforcement.

The Decoder1 day agoPolicy
Image: The Decoder

Anthropic recently disclosed a major safety lapse, revealing that its biological classifiers designed to block queries about chemical and biological weapons were inactive from May 2025 through April 2026. The nearly year-long outage left a massive volume of interactions unprotected. Specifically, the gap affected approximately 133 million chat sessions generated by a pool of roughly 50,000 external contractors who provide human feedback to train the company's models.

The safety filters are engineered to prevent users from extracting dangerous, actionable instructions regarding hazardous biological agents. During the outage, the external contractors interacting with the unfiltered models were vetted only by third-party vendors. Anthropic admitted that these external screening processes were frequently insufficient. Despite the prolonged exposure, the company stated that an internal investigation uncovered no evidence of actual misuse or malicious attempts to acquire weapon-related knowledge.

This security failure is particularly notable given the public stance of Anthropic's leadership. The company's chief executive has previously warned that the AI-assisted creation of chemical and biological weapons represents a far more severe threat than cyberattacks. In response to the incident, Anthropic has implemented stricter vetting requirements for its external contractors to prevent future vulnerabilities.

For AI practitioners and safety researchers, this incident underscores the operational challenges of maintaining robust guardrails during rapid development. It highlights how easily critical safety layers can fail silently in production environments. Separately, Anthropic recently adjusted its safety posture by loosening classifiers on its Fable 5 model. This decision followed complaints from researchers who argued that overly aggressive filters were hindering legitimate scientific inquiry, illustrating the delicate balance developers must strike between safety and utility.

This is our own summary of reporting by The Decoder

More in Policy