Data and Environment Curation for Post-training LLMs
AI Engineer World's Fair 2026 · 19:12
Reinforcement learning environments and agent optimization
Bespoke Labs builds reinforcement learning environments and infrastructure that help AI labs and enterprises train and evaluate agents for complex, multistep work. Its company-scale simulations incorporate real codebases and microservices, letting agents practice long workflows and enterprises test behavior in environments that mirror their systems. Its infrastructure combines environment creation, sandboxed execution, and automated prompt and policy optimization.
The company was co-founded by CEO Mahesh Sathiamoorthy, previously at Google and DeepMind, and Chief Scientist Alex Dimakis, a machine learning professor at UC Berkeley. Its collaborative research includes OpenThoughts, which supplies reasoning datasets and models alongside publicly available data generation, training, and evaluation code. Another contribution, GEPA, uses natural-language reflection on agent execution traces to diagnose problems, propose and test prompt changes, and combine improvements through evolutionary search. These projects address both the data used to train reasoning models and the instructions that guide agents.
Bespoke works with frontier labs on training infrastructure and with enterprises on optimizing production agents. As of August 2026, the company reported more than 200 teams using GEPA in production. In 2026, it announced $40 million raised across seed and Series A financing, with Wing VC leading the Series A.
AI Engineer World's Fair 2026 · 19:12
Affiliations reflect their AIE appearances, not necessarily current employment.
Start with Ryan Marten’s talk for an introduction to OpenThoughts3 and audience questions about supervised fine-tuning and reasoning-trace failures.
Ryan MartenAI Engineer World's Fair 2025
Continue with Mahesh Sathiamoorthy’s talk to follow question selection, filtering, and answer generation, and learn how these concerns connect to enterprise agent deployment.
Mahesh SathiamoorthyAI Engineer World's Fair 2026
Marten discusses how teacher selection, domain-specific filtering, and dataset scaling affect performance, with evaluation on AIME, LiveCodeBench, and GPQA Diamond.
Sathiamoorthy traces Bespoke Curator and Bespoke Stratos into OpenThoughts, then discusses sandbox infrastructure, agent benchmarks, and prompt-and-harness optimization.
Affiliations reflect each recorded session, not necessarily current employment.