← All organizations

AI computing hardware and cloud infrastructure

Cerebras

Cerebras builds processors, supercomputers and cloud services for AI training and inference. Its Wafer-Scale Engine technology powers systems that organizations use to build on-premises supercomputers for medical research, energy and agentic AI. Developers and enterprises can also access Cerebras through pay-as-you-go cloud offerings. Its announced CS-4 system combines three WSE-3 Turbo processors in a rack, with initial shipments scheduled for the third quarter of 2026.

Founded in 2015 by Andrew Feldman, Gary Lauterbach, Michael James, Sean Lie and Jean-Philippe Fricker, Cerebras remains led by Feldman as CEO. Its architecture connects compute cores across a silicon wafer and places local memory beside each core. Its training approach also separates model storage from computation: MemoryX streams weights to the processors, while SwarmX distributes weights and combines gradients across systems. This lets training capacity expand through data parallel replication without partitioning the model across devices.

Cerebras serves both scientific computing and commercial AI workflows. Customers including Block, Figma, AlphaSense and GSK use its inference services for agentic workflows, while Cognition and Lovable have signed cloud capacity agreements. The company completed its IPO in May 2026, raising approximately $6.38 billion in gross proceeds, and trades on Nasdaq under CBRS.

www.cerebras.ai

2 talks

Newest first

3 speakers at AIE

Affiliations reflect their AIE appearances, not necessarily current employment.

Start here

  1. Fast Models Need Slow Developers

    Start here to learn how memory bandwidth, on-chip SRAM, KV caches, and separated prefill/decode workloads contribute to inference speed.

    Sarah ChiengAI Engineer Europe 2026

  2. From Mixture of Experts to Mixture of Agents … with Super Fast Inference

    Choose this workshop for application-building steps: obtaining a Cerebras API key, deploying a GitHub-based Streamlit application, and configuring prompts.

    Daniel Kim · Daria SobolevaAI Engineer World's Fair 2025

Messages from the stage

Planning and execution at different speeds

Chieng recommends pairing larger planning models with fast execution models, actively steering agents, and keeping context in external progress files.

From model efficiency to model collaboration

Kim and Soboleva explain Mixture of Experts as an approach to efficient language-model scaling, then introduce Mixture of Agents as a collaborative multi-model inference architecture.

Affiliations reflect each recorded session, not necessarily current employment.

Company sources · checked 2026-08-28