Cerebras has introduced the fourth generation of its AI system: CS-4. Built from three new Wafer Scale Engine 3 Turbo processors with a completely redesigned rack, CS-4 delivers up to 30 times faster inference than GPU systems and a simple path to deploying hyperscale capacity.
Key Performance Claims
- Up to 30x faster than GPU systems — a new record for production inference
- 1,000+ tokens per second for models exceeding 10 trillion parameters, enabled by wafer-to-wafer interconnect latency as low as 2 microseconds
- Up to 10x more throughput per watt than CS-3, with up to 2x faster performance
- Native support for disaggregated inference — pairing purpose-built prefill engines with CS-4’s ultra-low-latency decoding
The system targets a broad audience: developers wanting responsive reasoning at 30x speed, data center operators wanting higher throughput per gigawatt, and neoclouds/hyperscalers needing modular systems that can be manufactured, installed, and upgraded at gigawatt scale.
The Nexus Platform Architecture
CS-4 is the first system built on Cerebras’ new Nexus rack-scale platform, which rethinks the rack around compute, power, and I/O:
- Wafer-Scale Backpack — a pluggable compute design that folds power conversion, liquid cooling, I/O, and control electronics into a self-contained package with 50% fewer components, reducing deployment time from days to hours
- High-density power delivery — moving power conversion 100x closer to the processors, delivering twice the power to the WSE-3 Turbo and enabling higher operating frequencies
- New I/O subsystem — doubles bandwidth with standards-based RoCE v2 RDMA plus Direct Wafer Links for switch-free connections between systems
Disaggregated Inference
A notable architectural choice: CS-4 natively supports splitting inference into prefill (processing the incoming prompt) and decode (generating the response). Operators can use GPU or ASIC infrastructure — including AMD Helios and AWS Trainium — for efficient prefill, while CS-4 handles ultrafast decoding. This heterogeneous approach lets providers combine infrastructure economics with Cerebras decode speed.
Availability
The first CS-4 shipments begin this quarter. The launch extends Cerebras’ momentum in the inference speed race, weeks after revealing how it serves GPT-5.6 Sol at up to 750 tokens per second in its OpenAI ultrafast partnership.