Cerebras Unveils CS-4 to Run AI 30 Times Faster Than GPUs
Cerebras has unveiled its CS-4 wafer-scale system and a multi-year roadmap, aiming to deliver up to 30 times faster AI inference than traditional GPU setups to accelerate agentic workloads.

At the Hot Chips 2026 conference, Cerebras detailed its new CS-4 architecture, scheduled for general availability in the third quarter of 2026. Built on the reusable Nexus rack-scale platform, the CS-4 packs three WSE-3 Turbo wafers per rack to deliver 750 petaflops of performance, with each chip contributing 250 petaflops. This represents a significant leap from the CS-3, which offered 125 petaflops on a single WSE-3 wafer. Cerebras claims the CS-4 delivers up to twice the token-generation speed of its predecessor and runs frontier models up to 30 times faster than competing GPU setups.
The system achieves these speeds through radical hardware engineering. Cerebras placed AC/DC power converters just 0.5 millimeters from the wafer—100 times closer than the 50-millimeter distance found on standard GPU boards. This layout eliminates printed circuit board resistance, allowing the 5-nanometer silicon to double its clock speed. Additionally, a single WSE-3 Turbo wafer provides 53.5 petabytes per second of aggregate on-wafer bandwidth. By comparison, an Nvidia NVL72 rack provides 260 terabytes per second of NVLink bandwidth across roughly 5,000 cables, representing a 200-fold advantage for Cerebras in scale-up bandwidth.
Looking ahead, the company's roadmap outlines the CS-5 for 2027 and the subsequent CS-6. The CS-5 targets 10,000 tokens per second per user on open-source models like Gemma 4 31B and gpt-oss-120b. For trillion-parameter models like Kimi and GPT-5.6 Sol, it aims for 5,000 tokens per second per user and 3 million tokens per second per megawatt, with the capacity to run models exceeding 50 trillion parameters. To overcome memory limitations, the CS-6 will integrate 3D-stacked DRAM directly on top of wafer-scale SRAM, expanding capacity while maintaining data locality.
For AI practitioners, these advancements shift the focus of hardware competition from raw throughput to per-user latency. Extremely high token speeds are critical for agentic workloads, which chain multiple sequential model calls together. By accelerating individual steps, a multi-step agent chain becomes vastly more responsive. Furthermore, keeping compute on a single giant wafer avoids the expensive communication overhead of routing mixture-of-experts models across thousands of networked GPUs.
This is our own summary of reporting by AlphaSignal



