Cerebras unveils CS-4, doubling CS-3 performance on the same chip

Cerebras unveils CS-4, doubling CS-3 performance on the same chip

Cerebras has introduced the CS-4, a new rack-scale AI accelerator system that CEO Andrew Feldman calls the fastest system in the industry. A rack-scale product is a full server cabinet, complete with compute units, power and cooling, that goes into a data center as a single unit. The CS-4 still runs on the same 5nm WSE-3 chip used in the earlier CS-3, but doubles that system's performance by boosting clock speed through more power and better cooling. A single CS-4 rack now holds three wafers, up from two in the previous generation, and delivers up to 4,400 tokens per second per user. Cerebras says that throughput is up to 30 times faster than setups running on Nvidia GPUs. Memory capacity per wafer stays unchanged at 44 GB. The system also introduces a new modular "Backpack" design meant to speed up assembly, along with disaggregated inference built through partners including AMD and AWS Trainium. Independent analysts at SemiAnalysis reviewed the announcement and characterized the CS-4's networking gains specifically as fairly small, a more measured read than Cerebras's own framing. Cerebras says more details will follow at the Hot Chips conference. Cerebras hardware already sees real-world use: OpenAI runs it for Codex Spark, among other deployments.

Key facts

  • The CS-4 doubles the performance of the CS-3 while keeping the same 5nm WSE-3 chip, gained through higher clock speed via more power and better cooling.
  • A single rack now holds three wafers instead of two, and the system delivers up to 4,400 tokens per second per user.
  • Cerebras claims the CS-4 is up to 30 times faster than comparable Nvidia GPU setups; memory stays at 44 GB per wafer.
  • SemiAnalysis analysts assess the networking improvements as fairly small, a more cautious take than Cerebras's marketing.
  • OpenAI already uses Cerebras hardware for Codex Spark; AMD and AWS Trainium are named as partners for disaggregated inference.

Why it matters

The CS-4 is a significant generational jump on paper: double the throughput of the CS-3 without a new chip, achieved purely through system-level engineering, more power, better cooling, and packing three wafers into a rack instead of two. At up to 4,400 tokens per second per user, Cerebras is pushing inference speed as its main competitive edge over GPU-based providers, and the claimed 30x advantage over Nvidia setups is central to that pitch.

Who it affects

OpenAI is named as an existing customer, running Cerebras hardware for Codex Spark. AMD and AWS Trainium are named as partners in a disaggregated inference setup that pairs Cerebras compute with other vendors' infrastructure. Any enterprise or developer choosing between GPU and wafer-scale inference providers is a downstream audience for these throughput claims, as is Nvidia, whose GPU-based inference speed is the explicit comparison point.

How to use it

No pricing, licensing terms or general availability timeline for the CS-4 is given. The article states only that more details are coming at the Hot Chips conference, and describes the new "Backpack" design as enabling faster assembly, without further technical detail on how it works.

How solid is it

The headline performance and speed claims, doubling the CS-3, up to 4,400 tokens per second, and up to 30 times faster than Nvidia, come from Cerebras itself, via CEO Andrew Feldman and the company's own framing. The one outside check in the piece is SemiAnalysis, whose analysts assessed the CS-4's networking gains as fairly small, a notably more restrained view than Cerebras's marketing on the rest of the system.

Risks and caveats

The core throughput and speed-versus-Nvidia figures are vendor-supplied and have not been independently benchmarked in the source material. SemiAnalysis's pushback on networking gains suggests not every part of the CS-4 story matches Cerebras's framing. No pricing, release date or availability window is disclosed, and the "Backpack" design is described only in terms of its stated benefit, faster assembly, with no technical specifics.