Research

Anthropic Says Claude Mythos and GLM-5.3 Can Hijack Code

Anthropic's Frontier Red Team revealed that Claude Mythos Preview and GLM-5.3 can execute control flow hijacks, marking a significant escalation in AI-driven cyber capabilities.

Simon Willison1 day agoResearch
Illustration generated for this story

A newly disclosed evaluation by Anthropic's Frontier Red Team has revealed that the latest generation of large language models has crossed a critical threshold in cybersecurity capabilities. In a series of automated tests, both Anthropic's own Claude Mythos Preview and the rival GLM-5.3 model demonstrated the ability to autonomously execute complex cyberattacks. This development highlights a rapid escalation in the offensive digital skills possessed by commercial artificial intelligence systems.

The red-teaming researchers evaluated several advanced models against a randomized selection of 100 tasks taken from an internal Binary Exploitation benchmark. During these trials, Claude Mythos Preview successfully developed full control flow hijacks in 6 percent of the tests. Meanwhile, the GLM-5.3 model achieved the same level of compromise in 4 percent of the trials. Control flow hijacking is a sophisticated exploitation technique where an attacker redirects the execution path of a program to run malicious code.

While these single-digit success rates may appear modest, they represent a stark departure from previous model generations. The Frontier Red Team noted that older systems, such as Claude Opus 4.6 and GLM-5.2, failed to succeed in a single trial under the same testing conditions. The sudden jump from zero capability to reliable exploitation indicates that AI developers have crossed what researchers described as a "meaningful threshold" in the spread of advanced cyber capabilities.

For cybersecurity practitioners and AI safety researchers, these findings signal a pressing need for stronger defensive guardrails. As models grow more adept at identifying and exploiting software vulnerabilities, the window for securing legacy codebases is rapidly closing. The fact that multiple independent model families are simultaneously acquiring these offensive capabilities suggests that automated binary exploitation will soon become a standard feature of frontier AI systems.

This is our own summary of reporting by Simon Willison

More in Research