Announcement • 9h
Cerebras Systems Unveils CS-4 AI Accelerator Cerebras Systems introduced the Cerebras CS-4, the fastest AI accelerator in the industry. The CS-4 is a rack-scale solution built from three new Wafer Scale Engines and revolutionary rack and system designs. The CS-4 is the first member of the next-generation Cerebras Nexus rack-scale platform architecture. It is up to twice as fast as the CS-3, bringing the CS-4s advantage in tokens-per-second-per-user over GPUs to up to 30x more. The CS-4 solution also delivers up to 10x more throughput per watt than the CS-3, vastly improving data center economics. The CS-4 delivers both higher-value tokens and more total tokens within a given power budget—enabling datacenters to be vastly more profitable. The rack scale CS-4 is built from three of the newly released Wafer Scale Engine 3 Turbo (WSE-3T). The CS-4 delivers 750 PFLOPs of AI compute, 7.2 terabits per second of I/O, and 129.6 petabytes per second of memory bandwidth. Total compute fabric bandwidth jumps to 160.5 petabytes per second, and wafer to wafer latency drops as low as two microseconds, enabling the creation of very large clusters and the support of models with over 50 trillion parameters. The CS-4 sets a new high watermark for inference speed. In a head-to-head comparison on GPT-OSS-120B, when given identical prompts, the CS-4 delivers more than 4,400 tokens second per user (TPS/user), up to 30 times faster than GPU solutions. CS-4 is powered by the newly announced WSE-3 Turbo (WSE-3T). Like the WSE-3, the WSE-3T is the largest AI processor ever built, containing four trillion transistors and 900,000 AI-optimized cores across 46,225 square millimeters of silicon, with 44GB of SRAM integrated directly on the wafer. The WSE-3T doubles AI compute to 250 PFLOPS per wafer and doubles memory bandwidth to 43.2 petabytes per second. The on-chip fabric bandwidth and off-chip I/O both double to 53.5 petabytes per second and 2.4 terabits per second respectively. I/O latency shrinks from five microseconds to as low as two microseconds. CS-4 is the first iteration of the new Cerebras Nexus Platform Architecture. It is built around a modular concept with three foundational elements: Compute, Power and I/O. Modularity enables each element to scale independently so innovations get to market faster. The modular architecture also supports extremely rapid deployment and upgrades. Cerebras has fundamentally re-imagined the compute subsystem into a rear mounted “backpack” that attaches vertically to the power array. Each Wafer-Scale Backpack is a self-contained assembly that folds power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact, three-dimensional package built directly around the wafer. By decoupling compute from the power supplies, the Wafer-Scale Backpack simplifies manufacturing and reduces deployment time from days to hours. Compared with the prior-generation system, the Wafer-Scale Backpack has 50% fewer components and uses 60% more automated manufacturing. The Nexus Platform Architecture drives significant power delivery improvements. By moving power conversion 100x closer to the processors – from roughly 50 millimeters away from the processor as on conventional GPU boards to approximately 0.5 millimeters – CS-4 nearly eliminates board-level power loss. This delivers twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation. CS-4 introduces a new programmable I/O subsystem that supports two connectivity modes while doubling I/O bandwidth and reducing latency. The fully programmable Wafer I/O Module supports standards-based RoCE v2 RDMA over Ethernet for seamless integration into existing infrastructure and with an ecosystem of heterogeneous systems. Aggregate off-wafer bandwidth doubles to 2.4 terabits per second per wafer and 7.2 terabits per CS-4 rack solution. The Wafer I/O Module also supports a new communication mode, called Direct Wafer Links, which enables wafers to be linked within and across racks without a switch. Direct Wafer Links enables wafer-to-wafer latency as low as two microseconds. This low latency communication will allow the creation of massive CS-4 clusters and the ability to support models with more than 50 trillion parameters. Programmable low latency I/O is particularly beneficial for heterogeneous disaggregated inference, in which a purpose-built prefill engine processes an incoming prompt and then hands it to Cerebras for ultra-low latency decode. The combination of programmability, standards-based interfaces, and very low latency will allow the rapid creation of disaggregated solutions from different Cerebras ecosystem partners, like AMD Helios and AWS Trainium. Metric comparison: CS-3 (one wafer) vs. CS-4 (3 wafers): AI compute: 125 PFLOPS vs. 750 PFLOPS; Memory bandwidth: 21.6 PByte/s vs. 129.6 PByte/s; On-chip fabric bandwidth: 26.7 PByte/s vs. 160.5 PByte/s; System I/O bandwidth: 1.2 Tbit/s vs. 7.2 Tbit/s; I/O latency: 5 microseconds vs. 2 microseconds. First CS-4 shipments begin this quarter. Full system specifications are available in the CS-4 datasheet. Major Estimate Revision • Aug 19
Consensus EPS estimates fall by 32%, revenue upgraded The consensus outlook for fiscal year 2026 has been updated. 2026 revenue forecast increased from US$864.9m to US$887.3m. Forecast EPS reduced from -US$3.30 to -US$4.36 per share. Semiconductor industry in the US expected to see average net income growth of 57% next year. Consensus price target broadly unchanged at US$292. Share price fell 6.3% to US$220 over the past week. Recent Insider Transactions Derivative • Aug 18
Lead Independent Director notifies of intention to sell stock Eric Vishria intends to sell 68k shares in the next 90 days after lodging an Intent To Sell Form on the 17th of August. If the sale is conducted around the recent share price of US$230, it would amount to US$16m. Eric currently holds less than 1% of total shares outstanding. Company insiders have collectively sold US$21m more than they bought, via options and on-market transactions in the last 12 months. Live News • Aug 15
Cerebras Powers OpenAI Ultrafast Mode for GPT-5.6 Sol With 14x Speed Boost Cerebras Systems announced it is powering “Ultrafast mode” for OpenAI’s GPT-5.6 Sol in the OpenAI API, a new service tier that runs up to 14x faster than standard processing and is in limited preview for customers.
The company’s wafer-scale engine removes typical GPU memory bandwidth bottlenecks, which is the hardware basis for Ultrafast mode’s higher throughput without adding latency for OpenAI users.
Cerebras Systems shares trade at US$218.98, with the stock down 29.6% year to date, reflecting recent pressure despite the new OpenAI integration and updated guidance commentary.
This OpenAI tie-up puts Cerebras hardware directly behind a high-profile production workload, which could help validate its architecture for large-scale inference. The main risk is execution. Investors will be watching how quickly limited preview moves to broader availability and how far this translates into recurring cloud service revenue versus one-off or concentrated deployment spend. Announcement • Aug 14
Cerebras Powers Ultrafast Mode For OpenAI’s GPT-5.6 Sol Cerebras powered OpenAI's most capable model at up to 14x the speed, giving frontier intelligence with no compromise on latency. Cerebras announced that it is powering Ultrafast mode, a new service tier in the OpenAI API for GPT-5.6 Sol. Available initially in limited preview to OpenAI customers, Ultrafast runs GPT-5.6 Sol at up to 750 output tokens per second and up to 14x faster than Standard processing. GPT-5.6 Sol Ultrafast powered by Cerebras runs with the same intelligence as GPT-5.6 Sol Standard, enabling frontier intelligence at Cerebras' blistering fast speed. Ultrafast powered by Cerebras offers access to the full intelligence of GPT-5.6 Sol at speeds suited to work that cannot wait. Based on output speeds for Anthropic models reported by Artificial Analysis, Ultrafast is 5x faster than Claude Opus 4.8 in Fast mode, and 11x faster than Claude Fable 5. Benchmarks focused on economically valuable work, from programming to drafting legal documents, show how faster token generation translates directly into higher productivity for users. On Humanity’s Last Exam, a 2,500-question benchmark spanning graduate-level chemistry, economics and literature, GPT-5.6 Sol Ultrafast answered the full question set in just over 11 hours. This compares to more than three days of continuous compute for Claude Fable 5, with GPT-5.6 Sol Ultrafast reaching comparable accuracy nearly 7x faster. On GDP-Val, a benchmark of economically valuable knowledge-work tasks such as legal briefs, financial models, and engineering reports, Ultrafast delivered a 5.6x end-to-end speedup with no loss in quality. Ultrafast’s speed comes from Cerebras' Wafer-Scale Engine architecture, which keeps model weights on-chip — 44 GB of SRAM on each wafer-sized chip — rather than shuttling them between on-chip memory and off-chip storage as GPU-based inference must. This eliminates the memory-bandwidth bottleneck that constrains frontier-model inference speed on conventional hardware. Reported Earnings • Aug 13
Second quarter 2026 earnings: Revenues exceed analysts expectations while EPS lags behind Second quarter 2026 results: US$2.98 loss per share (down from US$5.87 profit in 2Q 2025). Revenue: US$180.1m (up 74% from 2Q 2025). Net loss: US$450.5m (down 246% from profit in 2Q 2025). Revenue exceeded analyst estimates by 8.4%. Earnings per share (EPS) missed analyst estimates by 78%. Revenue is forecast to grow 45% p.a. on average during the next 3 years, compared to a 25% growth forecast for the Semiconductor industry in the US.