OpenAI Runs GPT-5.6 Sol at 750 Tokens per Second
OpenAI has previewed an Ultrafast API mode that runs its GPT-5.6 Sol model at 750 tokens per second using Cerebras hardware, drastically reducing latency for complex agentic workflows.

OpenAI has introduced a preview of its new Ultrafast mode, an API service tier that allows its proprietary GPT-5.6 Sol model to reach speeds of up to 750 tokens per second. This performance is powered by Cerebras hardware, specifically the Wafer-Scale Engine 3 (WSE-3), which bypasses traditional GPU memory bottlenecks by keeping model weights entirely local. The deployment is part of a massive, multi-year compute partnership between OpenAI and Cerebras valued at more than $20 billion.
The new speed tier represents a massive leap over existing benchmarks. At 750 tokens per second, GPT-5.6 Sol runs 14 times faster than standard processing and roughly five times faster than typical production GPU setups, which average about 150 tokens per second. It also marks a 7-to-10x improvement over its predecessor, GPT-5.5 XHigh, which operates at 70 to 100 tokens per second. While competitors like Groq achieve similar speeds, they only support open-source models, making OpenAI's offering the first frontier proprietary model to enter this speed class.
For developers and enterprise practitioners, this extreme throughput fundamentally changes how AI agents are designed. In complex workflows where an agent must chain 30 or 40 consecutive model calls to complete a single back-office task, the speed difference compounds. A process that previously took minutes on standard hardware can now finish in seconds. This makes the model highly viable for latency-sensitive applications like real-time voice, financial research, incident response, commerce, and live interactive experimentation.
Currently, OpenAI is limiting access to Ultrafast mode to a select group of API customers. The company plans to expand capacity gradually, allowing businesses to sign up for updates as more infrastructure becomes online.
This is our own summary of reporting by AlphaSignal


