Build enterprise generative AI apps using Llama-3 at 1,000 tokens/s on the SambaNova AI platform
Varun Badrinath Krishna · Petro Junior Milan · Rachelle Mattern
AI Engineer World's Fair 2024 · 54:34
AI inference chips, systems, and cloud services
SambaNova Systems builds AI inference chips, systems, and cloud services for enterprises, AI labs, and service providers. Its SambaRack systems run models in customer data centers, while OpenAI-compatible APIs let developers connect applications to inference. SambaOrchestrator manages model deployments, monitoring, and automatic scaling across data centers. Together, these products support running large models and switching among models for agent workflows.
Founded in 2017, SambaNova was co-founded by CEO Rodrigo Liang, chief technologist Kunle Olukotun, and Christopher Ré. Its engineering centers on Reconfigurable Dataflow Units, which combine dataflow processing with three-tier memory. Its 2024 SN40L research described how streaming dataflow and SRAM, HBM, and DDR memory work together to address model-serving memory constraints. The paper demonstrated Composition of Experts with 150 specialized models and a trillion total parameters, addressing the cost of hosting multiple models and switching between them.
In July 2026, SambaNova completed the first close of $1 billion in Series F financing, led by General Atlantic at an $11 billion post-money valuation. JPMorganChase selected its SN40 and SN50 systems for on-premises inference. That month, the company reported a $3.5 billion order from Vector Core Compute for deployment over three years.
Varun Badrinath Krishna · Petro Junior Milan · Rachelle Mattern
AI Engineer World's Fair 2024 · 54:34
Rachelle Mattern · Petro Milan · Varun Krishna
AI Engineer World's Fair 2024 · 1:00:58
Affiliations reflect their AIE appearances, not necessarily current employment.
Start with this workshop for the progression from Python and API configuration to a working document-question-answering pipeline.
Varun Badrinath Krishna · Petro Junior Milan · Rachelle MatternAI Engineer World's Fair 2024
Use this workshop for its coverage of LangChain prompt composition, SambaStudio inference and embedding APIs, and PDF loading and splitting.
Rachelle Mattern · Petro Milan · Varun KrishnaAI Engineer World's Fair 2024
The hands-on material follows document loading, embeddings, and ChromaDB vector indexing to assemble a document-question-answering workflow.
The presenters introduce Samba-1's Composition of Experts architecture alongside demonstrations of Llama 3 and Samba-1 Turbo. Inference speed provides context for the application-building exercises.
Affiliations reflect each recorded session, not necessarily current employment.