Cerebras announced CS-4, its fourth-generation server rack, powered by three WSE-3 Turbo chips and built around a new architecture called Nexus, with first shipments starting this quarter1.
The CS-4 doubles the inference performance of its predecessor, the CS-3, while using the same third-generation 5nm wafer-scale engine2. Cerebras achieves the performance gain not through a silicon node shrink but by doubling clock speeds, enabled by redesigned power delivery and cooling systems that push the rack to 125–135kW TDP. The result: CS-4 doubles tokens per second per user per wafer compared with CS-3, at roughly the same bill-of-materials cost per wafer.
SemiAnalysis expects CS-4 to hit near 4,000 tokens per second per user on frontier models, up from approximately 2,000 on CS-3. Cerebras markets the system as delivering up to 30x interactivity improvement compared with GPUs and 43 PB/s of total on-chip memory bandwidth. Each wafer retains 44GB of SRAM capacity.
The rack architecture has been substantially reworked. Cerebras split the CS-4 into a front half dedicated to power delivery and a rear half dedicated to compute, packaged as modular, pluggable "backpacks". Each backpack houses a single wafer-scale engine; a CS-4 rack holds three, up from two wafers per rack in CS-3. Cooling infrastructure — pumps and heat exchangers — has been moved out of the rack entirely. The backpack design allows customers to install power modules first and socket in the wafer backpack on site, improving deployment times.
Networking receives a parallel upgrade. Off-wafer I/O doubles to 2.4 Tb/s from 1.2 Tb/s on CS-3, and latency through the two-layer fat-tree network drops to 3 microseconds from 5 microseconds. A new Wafer I/O interface — an upgraded FPGA card functioning as a NIC — converts Cerebras's proprietary I/O to standard Ethernet and is field-upgradeable, letting Cerebras adopt new networking standards without a chassis redesign. Direct wafer-to-wafer links are now possible, bypassing the switched network, with configurable routing through the FPGA.
Cerebras said a new I/O module will enable open, heterogeneous, and disaggregated inference architectures. SemiAnalysis noted that disaggregated setups pairing CS-4 with HBM-based systems could help overcome the wafer's memory capacity constraints. For mixture-of-experts models, every expert for a given model will sit on a single wafer, interleaved.
Cerebras has stated a roadmap target of roughly 2x faster performance every year, with a specific goal of 20x throughput improvement by 2027. The company said it has already co-designed the next-generation wafer-scale engine alongside Nexus.
The announcement arrives days after OpenAI on August 13 previewed its Ultrafast API tier running GPT-5.6 Sol at up to 14x standard speed, powered by Cerebras hardware ctx. ANALYSIS The CS-4 launch gives Cerebras a concrete product story to pair with that high-profile deployment, offering inference customers a path to double per-wafer throughput without increasing hardware spend.