Models

Google Launches Gemini 4 Argon with 1M Token Output

Google DeepMind has introduced Gemini 4 Argon, a frontier AI model designed for complex reasoning tasks that features an unprecedented one-million-token output limit.

DeepMind Blog1 day agoModels
Image: DeepMind Blog

Google DeepMind has unveiled Gemini 4 Argon, its latest frontier artificial intelligence model engineered to handle intricate, long-horizon professional workflows. The model is initially rolling out to select cybersecurity defenders through Google's Fairwind Program before a wider release to paid API customers and Google AI Ultra subscribers. During an introductory period, Argon will cost $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached inputs. Afterward, the price will rise to $4 per million input tokens and $20 per million output tokens.

For developers, Argon's key advancement is its one-million-token output limit, up from 64,000 tokens, enabling deep, multi-step reasoning in a single run. Google has used Argon internally to optimize quantum computing subroutines, beating a published baseline by 40% in minutes. Argon agents also analyzed telemetry to save over 300 TiB of memory across Google data centers, with total savings estimated between 500 TiB and 1 PiB. Additionally, the model migrated C/C++ codebases to Rust, rewriting 32,000 lines of SIMD code for the libgav1 video decoder to make it memory-safe and 2.7 times faster.

The model establishes several new performance benchmarks. It achieves a state-of-the-art 77.9% on DeepSWE v1.1 for software engineering and ranks first on Zapier's AutomationBench with a score of 51.3%. For multimodal tasks, Argon scored 91.7% on the LVBench long video understanding benchmark. In specialized domains, it leads on the Vals Index, Vals Finance Agent v2, and Harvey's Legal Agent Benchmark. In cybersecurity, Argon tied for first place on CWE-bench v1 with a score of 68%, demonstrating significant improvements over the older 3.8 Flash Cyber model.

To empower cyber defenders, Google is releasing Argon to trusted partners without cybersecurity guardrails, allowing it to autonomously discover and patch vulnerabilities. In early testing, security firm Wiz used the model to uncover a critical vulnerability in global hospital software. To mitigate risks, Google is implementing defenses against prompt injections, where it leads on Gray Swan's Indirect Prompt Injection benchmark, and is monitoring the model's chain-of-thought to prevent misalignment.

This is our own summary of reporting by DeepMind Blog

More in Models