Cerebras has introduced the CS-4, a rack-scale AI inference system that doubles the performance of the earlier CS-3 without moving to a new processor generation. It still uses the 5-nanometer WSE-3 wafer-scale chip, but runs at a higher clock speed with increased power and improved cooling.
A rack now contains three wafers instead of two and is rated for as many as 4,400 generated tokens per second for each user. Cerebras says that can be up to 30 times faster than Nvidia GPU configurations. Memory remains 44 GB per wafer, so the upgrade emphasizes compute and system packaging rather than larger on-chip capacity. A modular Backpack design is intended to simplify assembly.
The company also describes disaggregated inference that can pair its hardware with AMD or AWS Trainium systems. The headline speed figures are vendor claims, and analysts cited by The Decoder view the networking improvements as modest. More architecture and benchmark details are due at the Hot Chips conference, where independent comparisons should clarify model choice, batch size, latency and cost.