- Sources: Cerebras CS-4 page, HN discussion
- Summary: Cerebras states each CS-4 system carries three WSE-3 Turbo wafers, that power delivery sits 0.5 millimeters from the processor against roughly 50 millimeters on conventional GPU boards, and that a new programmable I/O subsystem links wafers within and across racks without a switch at wafer-to-wafer latency as low as two microseconds. The claimed results are up to 30 times faster inference than production GPU systems, up to 10 times more throughput per watt than CS-3, and more than 1,000 tokens per second on models above 10 trillion parameters. The performance figures are vendor claims with no independent benchmark located, and Cerebras' own footnote says comparisons rest on third-party benchmarking or internal testing and may vary by workload and configuration.
- Why it matters: Cerebras states first shipments begin this quarter, so the claims become checkable against customer workloads within months.
- Follow-up: Track any independent benchmark of the CS-4 against the stated GPU baselines.
send feedback on this story