AI computing hardware and cloud infrastructure
Cerebras
Cerebras builds processors, supercomputers and cloud services for AI training and inference. Its Wafer-Scale Engine technology powers systems that organizations use to build on-premises supercomputers for medical research, energy and agentic AI. Developers and enterprises can also access Cerebras through pay-as-you-go cloud offerings. Its announced CS-4 system combines three WSE-3 Turbo processors in a rack, with initial shipments scheduled for the third quarter of 2026.
Founded in 2015 by Andrew Feldman, Gary Lauterbach, Michael James, Sean Lie and Jean-Philippe Fricker, Cerebras remains led by Feldman as CEO. Its architecture connects compute cores across a silicon wafer and places local memory beside each core. Its training approach also separates model storage from computation: MemoryX streams weights to the processors, while SwarmX distributes weights and combines gradients across systems. This lets training capacity expand through data parallel replication without partitioning the model across devices.
Cerebras serves both scientific computing and commercial AI workflows. Customers including Block, Figma, AlphaSense and GSK use its inference services for agentic workflows, while Cognition and Lovable have signed cloud capacity agreements. The company completed its IPO in May 2026, raising approximately $6.38 billion in gross proceeds, and trades on Nasdaq under CBRS.
2 talks
Newest firstFrom Mixture of Experts to Mixture of Agents … with Super Fast Inference
AI Engineer World's Fair 2025 · 53:15
3 speakers at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
Start here
- Fast Models Need Slow Developers
Start here to learn how memory bandwidth, on-chip SRAM, KV caches, and separated prefill/decode workloads contribute to inference speed.
Sarah ChiengAI Engineer Europe 2026
- From Mixture of Experts to Mixture of Agents … with Super Fast Inference
Choose this workshop for application-building steps: obtaining a Cerebras API key, deploying a GitHub-based Streamlit application, and configuring prompts.
Daniel Kim · Daria SobolevaAI Engineer World's Fair 2025
Messages from the stage
Planning and execution at different speeds
Chieng recommends pairing larger planning models with fast execution models, actively steering agents, and keeping context in external progress files.
From model efficiency to model collaboration
Kim and Soboleva explain Mixture of Experts as an approach to efficient language-model scaling, then introduce Mixture of Agents as a collaborative multi-model inference architecture.
Affiliations reflect each recorded session, not necessarily current employment.
Company sources · checked 2026-08-28
- Cerebras — Company History and Team
- Cerebras — Press Releases
- Cerebras Systems Announces Closing of Initial Public Offering
- Cerebras Systems Fast Inference Cloud Business Nearly Quadruples in Second Quarter 2026
- Cerebras Unveils CS-4: Up to 30 Times Faster than GPU-based Solutions
- Cerebras Architecture Deep Dive: First Look Inside the HW/SW Co-Design for Deep Learning [Updated]

