Agents

Cursor Launches Claude Fable 5.1 for Self-Verifying Code

Cursor has integrated Anthropic’s new Claude Fable 5.1 model, bringing advanced self-verification capabilities that allow developers to run complex, hours-long coding tasks unattended.

AlphaSignal12 hrs agoAgents
Image: AlphaSignal

Cursor has enabled Anthropic's Claude Fable 5.1 directly within its code editor, introducing what its team calls the strongest model they have ever evaluated. Running at maximum effort on the multi-file CursorBench 3.2.0, Fable 5.1 achieved a score of 73.4%. This performance edges out its predecessors, Fable 5 and Opus 5, which scored 70.5% and 70.0% respectively on the same evaluation. The update introduces a thinking variant and a 300K context window to the editor's model selection menu.

The model shows massive improvements in agentic capabilities across several other industry benchmarks. On Terminal-Bench 4.0, Fable 5.1 scored 55.8%, which rises to 60.9% when utilizing the Mythos 5.1 safeguards profile, compared to 42.0% for Fable 5 and 52.3% for Opus 5. On Terminal-Bench-Science 0.1, it reached 52.6%, more than doubling Fable 5's 24.7%. Additionally, it scored 31.4% on AutomationBench, nearly doubling Fable 5's 17.1%. In browser-agent testing by Browserbase, Fable 5.1 completed 82% of tasks on their most difficult benchmark, surpassing Opus 5 at 74% and Fable 5 at 57% while consuming fewer tokens.

For software engineers, these numbers translate into a model that excels at self-verification. Rather than declaring a task complete while leaving broken tests behind, Fable 5.1 actively checks its own assumptions during long, unattended runs. This makes it ideal for open-ended, multi-hour agent tasks, greenfield prototyping, complex migrations, and root-cause debugging where cheaper models like Sonnet typically stall.

To access the model, users with Privacy Mode enabled must first accept the Fable 5.1 data retention policy in their Cursor Dashboard. Under Anthropic's policy, retained data is automatically deleted after 30 days, is not human-readable by default, and is not used for model training. Enterprise users seeking stricter controls can look to the upcoming Enterprise Frontier Safeguards program to keep data on their own cloud infrastructure.

This is our own summary of reporting by AlphaSignal

More in Agents