GPT-6 Astra Tops Stanford's New Science Benchmark
OpenAI's GPT-6 Astra topped Stanford's new Terminal-Bench-Science benchmark with a 63.3% success rate, exposing a massive gap between frontier models and open-weight alternatives.

Stanford researchers, alongside the Terminal-Bench and Harbor teams, have launched Terminal-Bench-Science 0.1.0, a new agentic benchmark designed to test frontier AI models on complex, multi-step scientific workflows. OpenAI's GPT-6 Astra, running at maximum reasoning effort, led the leaderboard with a 63.3% pass rate. Anthropic's Claude Opus 5.5 followed closely, achieving 61.9% at extra-high effort. The benchmark evaluates models using 70 expert-curated tasks across five domains, placing agents in a sandboxed terminal with real data, software, and up to eight hours (28,800 seconds) of execution time.
The evaluation revealed a massive performance gap between proprietary frontier models and open-weight alternatives. The leading open-weight models, GLM-5.3 and DeepSeek V4.1 Flash, scored near 10%, lagging more than 50 percentage points behind the leaders. Alibaba's Qwen3.8 Max reached 12%, while Fable 5.1 trailed the top two models by approximately 20 percentage points. The benchmark's all-or-nothing grading system, which requires passing every pytest-based check to succeed, highlights how easily minor formatting or calculation errors can derail automated scientific research.
The benchmark demonstrated that scaling up reasoning budgets can yield dramatic performance improvements, though not always linearly. Claude Opus 5.5 improved from 24% at low effort to 61.9% at extra-high effort, though this came with a fivefold increase in cost. However, at maximum effort, its score actually dropped to 59.0%. Similarly, GPT-6 Sol gained 27 percentage points between low and maximum effort, but at 7.5 times the cost. Across domains, life sciences proved the most challenging, while mathematics was the strongest. For instance, Claude Opus 5.5 at extra-high effort passed 71% of mathematical tasks but only 46% of life science tasks.
Unlike traditional evaluations that rely on multiple-choice questions, Terminal-Bench-Science requires models to execute end-to-end computational workflows, such as tracking cell divisions in mouse myoblasts or calculating asteroid trajectories. The benchmark's creators are currently accepting task submissions for the 0.2 release, with a deadline set for October 5.
This is our own summary of reporting by AlphaSignal



