Hardware

Cerebras Unveils CS-4 to Run AI 30 Times Faster Than GPUs

Cerebras has unveiled its CS-4 wafer-scale system and a multi-year roadmap, aiming to deliver up to 30 times faster AI inference than traditional GPU setups to accelerate agentic workloads.

AlphaSignal9 hrs agoHardware
Image: AlphaSignal

At the Hot Chips 2026 conference, Cerebras detailed its new CS-4 architecture, scheduled for general availability in the third quarter of 2026. Built on the reusable Nexus rack-scale platform, the CS-4 packs three WSE-3 Turbo wafers per rack to deliver 750 petaflops of performance, with each chip contributing 250 petaflops. This represents a significant leap from the CS-3, which offered 125 petaflops on a single WSE-3 wafer. Cerebras claims the CS-4 delivers up to twice the token-generation speed of its predecessor and runs frontier models up to 30 times faster than competing GPU setups.

The system achieves these speeds through radical hardware engineering. Cerebras placed AC/DC power converters just 0.5 millimeters from the wafer—100 times closer than the 50-millimeter distance found on standard GPU boards. This layout eliminates printed circuit board resistance, allowing the 5-nanometer silicon to double its clock speed. Additionally, a single WSE-3 Turbo wafer provides 53.5 petabytes per second of aggregate on-wafer bandwidth. By comparison, an Nvidia NVL72 rack provides 260 terabytes per second of NVLink bandwidth across roughly 5,000 cables, representing a 200-fold advantage for Cerebras in scale-up bandwidth.

Looking ahead, the company's roadmap outlines the CS-5 for 2027 and the subsequent CS-6. The CS-5 targets 10,000 tokens per second per user on open-source models like Gemma 4 31B and gpt-oss-120b. For trillion-parameter models like Kimi and GPT-5.6 Sol, it aims for 5,000 tokens per second per user and 3 million tokens per second per megawatt, with the capacity to run models exceeding 50 trillion parameters. To overcome memory limitations, the CS-6 will integrate 3D-stacked DRAM directly on top of wafer-scale SRAM, expanding capacity while maintaining data locality.

For AI practitioners, these advancements shift the focus of hardware competition from raw throughput to per-user latency. Extremely high token speeds are critical for agentic workloads, which chain multiple sequential model calls together. By accelerating individual steps, a multi-step agent chain becomes vastly more responsive. Furthermore, keeping compute on a single giant wafer avoids the expensive communication overhead of routing mixture-of-experts models across thousands of networked GPUs.

This is our own summary of reporting by AlphaSignal

More in Hardware