Customized, production ready inference with open source models: Dmytro (Dima) Dzhulgakov
AI Engineer World's Fair 2024 · 18:55
AI model training and inference infrastructure
Fireworks AI provides infrastructure for developers to train, customize, and serve AI models using their own data. Its platform connects training—including custom reinforcement learning loops—with production inference through serverless APIs, dedicated deployments, and reserved capacity. Customers can run existing open models or deploy their trained versions; Cursor uses Fireworks for coding models, while Harvey uses it for legal AI. Fireworks Nexus adds routing between open and closed models for coding workloads.
Founded in 2022, Fireworks has seven cofounders: CEO Lin Qiao, Benny Chen, Chenyu Zhao, Dmytro Dzhulgakov, Dmytro Ivchenko, James Reed, and Pawel Garbacki. Their backgrounds span Meta’s PyTorch and machine learning infrastructure and Google Vertex AI; Qiao previously headed PyTorch at Meta. Its FireOptimizer adapts speculative decoding to a customer’s workload: smaller draft models propose tokens for the main model to verify, with automated draft-model training and evaluation intended to reduce inference latency. In 2026, Fireworks acquired Hathora, adding global compute orchestration technology to its inference infrastructure.
In July 2026, Fireworks announced a $1.505 billion Series D at a $17.5 billion valuation. The company reported surpassing $1 billion in annualized revenue run rate and serving more than 40 trillion tokens per day. More than 95% of that volume came from models specialized on customers’ proprietary data and optimized for specific jobs.
AI Engineer World's Fair 2024 · 18:55
Lin Qiao · Dmytro (Dima) Dzhulgakov
AI Engineer World's Fair 2024 · 18:55
Affiliations reflect their AIE appearances, not necessarily current employment.
The material connects production inference to GPU deployment challenges, custom CUDA kernels, and optimization of long-prompt RAG through caching.
Function calling and structured JSON generation bring attention to how customized models connect with application workflows.
Affiliations reflect each recorded session, not necessarily current employment. Both listings describe overlapping material and name Dmytro (Dima) Dzhulgakov as substituting for scheduled speaker Lin Qiao.