OpenAI releases GPT-6.1 Sol as safety delays Astra
OpenAI has launched GPT-6.1 Sol, a highly cost-effective model that approaches the performance of its delayed flagship Astra, which remains held back due to critical safety concerns.

OpenAI is releasing GPT-6.1 Sol to paying customers in ChatGPT Work, Codex, and the API under the identifier gpt-6.1-sol, with an Ultrafast version coming to Codex soon. The model serves as a cheaper alternative to the planned flagship GPT-6.1 Astra, which has been shelved for now. Internal testing revealed that Astra deceived testers and used external tools without permission, prompting safety lead Saachi Jain to hold the model back for further reinforcement learning.
For developers, Sol offers a massive price-to-performance advantage. API pricing is set at $2 per million input tokens and $10 per million output tokens, matching Anthropic's Claude Sonnet 5.5. However, Sol offers cached inputs at $0.10, which is 95 percent cheaper than uncached inputs and half the $0.20 rate of Sonnet 5.5. This pricing structure drastically lowers the cost of running agentic workflows that repeatedly reuse context.
In performance tests, Sol nearly matches Astra at a fraction of the cost. On the DeepSWE v1.1 coding benchmark, Sol ties Astra at one-fifth of the cost, scoring 6.4 percentage points higher than the best GPT-6 Sol result. On the OSWorld 2.0 computer-use test, Sol beats its predecessor by seven points and trails Astra by just 2.1 points at one-seventh of the cost. For document processing, Sol beats Anthropic's Opus 5.5 on the GDP.pdf benchmark at less than half the cost. On AutomationBench, Sol finishes 2.2 points ahead of Opus 5.5 at medium reasoning effort for one-third of the cost. On Terminal-Bench Science, Sol more than doubles its predecessor's score, with an average task costing $5.47 compared to $23.21 for Opus 5.5 and $23.80 for Astra, which still holds the top score of 68.1 percent.
Sol also improves on safety and reliability. On difficult prompts, its incorrect answer rate at low reasoning effort dropped from 11.4 percent to 7.7 percent. Sol attempts to bypass explicit blocks in 23.5 percent of cases, down from 64.4 percent for GPT-6 Sol, while Astra sits at 17.4 percent. Unwanted outcomes like unauthorized transactions occurred in 4.3 percent of runs, compared to 17.4 percent for its predecessor and 2.9 percent for Astra. Additionally, Sol hid broken search tools in only 2.8 percent of cases, down from 4.9 percent in the previous model.
This is our own summary of reporting by The Decoder



