Models

Anthropic Debuts Claude Opus 4.6 with 1M Token Context

Anthropic has released Claude Opus 4.6, a powerful new model featuring a one-million-token context window and advanced reasoning designed to handle complex, multi-step developer workflows.

Anthropic13 hrs agoModels
Image: Anthropic

Anthropic has launched Claude Opus 4.6, its most powerful model to date, introducing a one-million-token context window in beta on the Claude Developer Platform. The model is designed to sustain longer agentic tasks and operate reliably across large codebases. To help developers manage long conversations, Anthropic introduced context compaction in beta, which automatically summarizes older context, alongside a massive 128k output token limit. Pricing remains at $5 per million input tokens and $25 per million output tokens, though premium rates of $10 and $37.50 apply for prompts exceeding 200k tokens. US-only inference is also available at 1.1 times the standard token pricing.

The model establishes new performance baselines across several industry evaluations. On the GDPval-AA knowledge work benchmark, Opus 4.6 outperforms OpenAI's GPT-5.2 by roughly 144 Elo points and its predecessor, Claude Opus 4.5, by 190 points. It achieved a 90.2% score on the BigLaw Bench, with 40% perfect scores and 84% of scores above 0.8, while earning $3,050.53 more than Opus 4.5 on Vending-Bench 2. On the MRCR v2 eight-needle needle-in-a-haystack test, Opus 4.6 scored 76%, vastly outperforming Claude Sonnet 4.5's 18.5%. It also leads on Humanity's Last Exam with a 53.0% score using tools, tops Terminal-Bench 2.0, and scored 81.42% on SWE-bench Verified with a prompt modification. In a Box evaluation, the model achieved a 68% score compared to a 58% baseline, and scored 62.7% on MCP Atlas at high effort.

For practitioners, the update introduces adaptive thinking, allowing the model to autonomously decide when to use extended reasoning. Developers can fine-tune this behavior using four effort levels: low, medium, high, and max. In Claude Code, developers can now deploy parallel agent teams to collaborate on tasks like codebase reviews. Additionally, Anthropic has upgraded Claude in Excel to handle unstructured data and released a research preview of Claude in PowerPoint. To address safety concerns, particularly regarding the model's advanced coding capabilities, Anthropic implemented six new cybersecurity probes to detect potential misuse.

This is our own summary of reporting by Anthropic

More in Models