Models

Google Gemini 3.8 Flash Outperforms Claude Opus 5

Google DeepMind has launched Gemini 3.8 Flash, a highly efficient model that outperforms larger rivals like Claude Opus 5 on key agentic benchmarks at a fraction of the cost.

AlphaSignal1 day agoModels
Image: AlphaSignal

Google DeepMind has launched Gemini 3.8 Flash, marking its third update to the mid-tier Flash lineup in roughly three months. The new model is immediately available to Gemini app subscribers on Pro and Ultra plans, as well as through Google AI Studio, the Gemini API, and Google Antigravity. This release follows Gemini 3.6 Flash in July and Gemini 3.7 Flash in mid-August, continuing Google's rapid cycle of performance upgrades without raising prices.

The model is priced at an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, after which it will rise to $1.50 input and $7.50 output. This pricing makes Gemini 3.8 Flash roughly six to seven times cheaper than Claude Opus 5, which costs $5 for inputs and $25 for outputs per million tokens, and significantly less expensive than GPT-5.6 Sol at $4 input and $20 output. Despite the lower cost, the new model beats both competitors on several key benchmarks.

In domain-specific agent tasks, Gemini 3.8 Flash achieved a 10.0% pass rate on the Vals Finance Agent V2 evaluation, surpassing GPT-5.6 Sol's 2.5% and Claude Opus 5's 6.7%. It also outperformed both on the Harvey Legal Agent Benchmark and edged out Opus 5 on Terminal-bench 2.1 with a score of 89.4% compared to 89.1%. Compared to its predecessor, Gemini 3.7 Flash, the new version showed significant improvements, jumping from 11.2% to 19.1% on Terminal-bench 4.0, from 43.5% to 56.5% on BioMysteryBench Human Difficult, and from 50.6% to 59.0% on OSWorld-2.0. Additionally, it scored 54.9% on the HLE-Verified expert reasoning evaluation and excelled on the DeepSWE v1.1 software engineering benchmark.

For developers running high-volume autonomous agents, long-horizon coding loops, or large-scale document processing, Gemini 3.8 Flash offers a highly cost-effective alternative to frontier models. However, Google's model card notes a slight regression in safety performance for non-English languages compared to Gemini 3.7 Flash. Practitioners should run their own regression tests on non-English pipelines before deploying the new model in production.

This is our own summary of reporting by AlphaSignal

More in Models