AI Engineer World's Fair 2026
Data and Environment Curation for Post-training LLMs
About this talk
Bespoke Labs CEO Mahesh Sathiamoorthy explains why high-quality synthetic data and reinforcement-learning environments are the main bottlenecks in post-training reliable, autonomous LLM agents. He traces Bespoke Curator and Bespoke Stratos into the collaborative OpenThoughts reasoning-data project, describes question selection, filtering, answer generation, and scaling, and discusses agent benchmarks, enterprise deployment, sandbox infrastructure, and prompt-and-harness optimization.
Chapters
- 0:00Introduction: Bespoke Labs, Curator, Stratos, and OpenThoughts
- 1:56Agent benchmarks and the post-training data bottleneck
- 6:06OpenThoughts collaboration and reasoning-data curation recipes
- 13:46Enterprise production deployment
- 17:27Sandbox infrastructure, prompt optimization, and agent-training stack
Talk transcript
- 0:00
[on-hold jingle] Hey, everyone. Um, today I'll be talking about data c-- and, uh, environment curation for, uh, post-training LLMs.
- 0:21
And I am Mahesh Sathiamoorthy. Um, I'm co-founder and CEO at Bespoke Labs, and previously I was a researcher and, uh, engineer at, uh, Google DeepMind. So very briefly, I will tell you a little bit about, uh, Bespoke, and, uh, after that, the talk will be mostly around, uh, open source work we have done.
- 0:41
So Bespoke is an applied data research lab with a mission to help enterprises and frontier labs access high-quality data and RL environments for their post-training needs. So very briefly, what we do and what we have done is that last year we put out something called Curator, which is a tool for curating, uh, synthetic data for post-training with
- 1:05
basically SFT. And right after that, actually DeepSeek landed, and we started an effort to curate reasoning data, and that's how we started something called Bespoke Stratos, which eventually formed some-- uh, into the project called OpenThoughts, which some of you hopefully know about.
- 1:22
And we have also been core contributors to Terminal-Bench. Um, you know, these days we, we do a lot of research and build and ship RL environments, so I was actually looking forward to the previous talk, uh, from Nick, who's also, you know, uh, doing something similar.
- 1:40
Uh, and, and the other thing we do is we do a lot of post-training and help enterprises, uh, to get their own custom models, right? That's the name of-- That's how we ended up with the Bespoke, uh, ta-title for the company.
- 1:56
Uh, the, the other thing I want to kind of mention is this, there is this, you know, you know, in, in our industry, there are a lot of people who create data, uh, create RL environments, and then there are the, uh, researchers who consume this.
- 2:09
But I feel like there is this slight mismatch, and it's kind of beneficial for someone to kind of go do both at the same time. And in fact, uh, as you're curating data, you want to put yourself in the shoes of the researcher to see what, what does it take to, you know, uh, actually move the metrics
- 2:28
on the models. So that's one of the motivations of how we kind of think about... The other thing I want to kind of talk about is, you know, how, you know, uh, AI has evolved, right?
- 2:40
So early on, we used to think about and evaluate models on what they know. Um, for example, this is a-- this was a very popular benchmark on, uh, testing LLMs on various kinds of STEM, humanities, and all that knowledge.
- 2:56
And these days, we have all these benchmarks that test, uh, how, how, how agents are able to do things. We have moved on from knowing to doing, right? So that's the idea of agents, obviously.
- 3:08
And the-- One of the key principles or one of the key things about agents is that they are autonomous. And there are, as I was saying, there are many benchmarks, including, uh, SWE-bench, Terminal-Bench, and so on.
- 3:21
But ultimately, for many people, what they care about is, are these agents autonomous for long durations of time? Uh, Nick had, uh-- Sorry, uh, Ross had a great talk on long horizon, right?
- 3:33
So that's the goal, is eventually we make these agents autonomous for maybe few hours or a, you know, few days or a few weeks. And what is it that's blocking the, uh, autonomy of agents?
- 3:48
It's basically reliability, right? So at some point, something falls apart. Like, either they call the wrong tool, or they made a mistake and, you know, what, whatnot, right? And what, what's one lever to improve reliability?
- 4:02
There, there are, of course, many. Um, obviously, you can prompt your way to improving the agent's, uh, reliability, or you can, uh, update the harness, you know, the tools and whatnot.
- 4:14
But post-training is a very powerful tool to improve reliability or, or maybe even pre-train good models, right? So if you think of frontier labs, this is one of their, uh, primary mechanisms of improving agents over to, to, uh, get [clears throat] better, uh, capabilities in,
- 4:33
um, va-various domains or, you know, uh, for, for better, uh, benchmark numbers or better, uh, um, autonomy for long-longer and longer durations. And for post-training, one of the popular techniques, as you know, is reinforcement learning, and that's kind of, um, something, you know, a lot of you are, um, excited about is the, uh, you know, notion of
- 4:58
RL environments. But ultimately, for post-training, be it SFT or, or, uh, reinforcement learning, data is the bottleneck, right? So when, when I talk about data, R-RLMs are also something I'm calling it as data.
- 5:11
It's just the data is now in a very different shape. Um, again, here, you know, compute is kind of well-defined. Models, you know, uh, good sort of models exist and the, uh, infrastructure to post-train, for example, uh, there, there are various providers like Fireworks, Tinker or, uh, SLIM-VIL and whatnot.
- 5:36
So RL, all of those are somewhat well-defined. Most of the places where people struggle, at-- especially enterprises, is that they don't have access to good quality data and RLMs.
- 5:47
And this obviously also applies to frontier labs, where they have all this infra set up and they are, you know, needing good quality RLMs, right? Uh, beyond-- So that, that's one of the-- This is kind of how we are thinking about at why to invest time in, you know, doing data research and RL and research.
- 6:06
And as a side note, one of the, um, other, uh, the-- there are many other benefits of post-training. For example, you can reduce latency or improve cost, throughput and whatnot, and I'll give one concrete example of a post-training work we did, uh, with one of the enterprises.
- 6:25
So in this talk, I will mostly, uh, cover some of the work we have done in the open source, uh, community. So we did some work on curating reasoning data for reasoning models and for, uh, you know, curating trajectories and, uh, en-environments for agents.
- 6:43
And recently, we had an engagement with post-training, which, uh, I'll, I'll very briefly talk about, and some tools on data curation.
- 6:51
So OpenThoughts, um, is, uh, reasoning dataset as well as a paper, right? So we, we started this effort last year. As I was saying, this, uh... We, we-- After DeepSeek came out, we realized that there is a lack of very high-quality reasoning data in the, uh, community.
- 7:12
Obviously, the labs have access to good data, but outside we didn't have access to data, right? So we, we at Bespoke started this effort called Bespoke Stratos, and then we realized that this is actually quite useful.
- 7:25
So we joined, um, together with various folks in, uh, Stanford, UC Berkeley, UW and so on to create this consortium called OpenThoughts, and we did a lot of work on basically identifying the curation recipe, and we also published this as a paper in ICLR of this year, and this is the main figure of the paper.
- 7:46
So what it shows is, like, we, we figured out a curation recipe, and it shows the scaling law, right? So again, this is last year when Amy, uh, and, and, and, uh, LiveCodeBench and these, these were some of the popular benchmarks.
- 8:01
What we showed is that with this recipe, if you keep, you know, scaling up the dataset size, the, the, you know, the, the... It's a scalable recipe, right? The, the metrics also improve.
- 8:12
Um, it's actually very widely used as well. For example, this is, um, uh, Microsoft CSO tweeting about the work, and this Ale-Alex is my, uh, co-founder. He's, uh, chief scientist and also a professor at UC Berkeley.
- 8:27
And this is John Schulman talking about OpenThoughts that he has-- he and his, uh, colleagues have been using it internally at, uh, Thinking Machines, right? And some of their blog posts also reference this.
- 8:39
So I'll ve-- uh, talk about how we did the curation for OpenThoughts. Um, th-this is the pipeline that we used. So you start with curate-- You start with a bunch of source questions, right?
- 8:52
So there are various datasets out there that have the, uh, prompt response, and we ch-choose with the prompts. We start with the prompts. These are various, uh, sources we have.
- 9:03
And then, uh, if you look at the paper... So if you look at this graph, for any given data point, say if there are ten thousand samples that you want, the question is then how do you choose, uh, the questions from all these different dataset so that you have ten thousand, uh, for the data point?
- 9:24
So the, the-- then there is the aspect around how do you mix these questions. So, uh, you can use various methods. So the paper talks about, for example, using LLMs to check for whether this is a good, good question, hardness of a question, and so on.
- 9:40
And then you want to filter questions, uh, and generate the answers. Again, this is all, like, driven by LLMs, right? So this is the curation recipe we did for, uh, creating this reasoning dataset.
- 9:52
And j-- The, the answer generation is using teacher models, so you can take other reasoning data, uh, uh, reasoning models such as DeepSeek or Qwen-based models, uh, or even Gemini and whatnot.
- 10:03
And then you can also filter the answers once you have the answers for, uh, these questions. Um, and, and then you can also, you know, given a question, generate multiple answers or a single answer.
- 10:16
So these are various knobs in the curation recipe, and the systematic way of doing this is, like, you run ablations and figure out which, uh, you know, in each of these stages what works, and you kind of proceed to the next.
- 10:31
So after doing all of this, you get the final recipe, right? So this, uh, you can read this paper. It has lots and lots of, uh, you know, information about how we did the curation.
- 10:42
But here are some of the learnings that, you know, some of them are quite, uh, counterintuitive, and some of this was also covered in last year's, uh, AI, uh, AI engineer conference.
- 10:54
For example, sampling, um, multiple answers per question works pretty well. This is something that, uh, we... It's, it's kind of counterintuitive. So as an example, something else we could have done is we could have had more, much, many more questions and then just answered them e-exactly once versus taking one question and answering them sixteen times.
- 11:18
The-- I think the, the reasoning is probably that it gives like a variety of how reasoning is done. So the, the-- During fine-tuning, we also use the, the reasoning traces, right?
- 11:29
So I think the diversity helps there. And the other thing we saw is, like, the stronger teachers are not always the best, uh, uh... St-stronger models are not always the better teachers.
- 11:42
And there were a few other counterintuitive aspects around like, you know, uh, synthetic question generation or qu-question answering working, whereas answer, answer filtering and other aspects not working very well.
- 11:54
And after the OpenThoughts work, which was around, uh, data curation for reasoning models such as, you know, uh, DeepSeek kind of models, we moved on to OpenThoughts Agents, which is, um, very similar, but how do you curate these, the, the data and RL environments for, uh, training agents now, right?
- 12:13
Not reasoning models. We, we have a very similar figure here. Again, we want to establish scaling loss. Um, so as you increase the dataset size, we want to make sure that the curation recipe actually works.
- 12:25
Uh, and again, I'm, I'm not going to go into details here, but very similarly, there are various ways of choosing different sources, for example, Stack Exchange and, and whatnot.
- 12:38
How do you mix the tasks? How do you filter generating the rollouts, fil-- uh, choosing the teacher, and so on? And again, th-these are some of the lessons, learnings.
- 12:48
Um, a-as an example, e-ev-even here we saw that stronger models are not necessarily the, uh, best teachers, right? So we found out some, some, some of the, I think, uh, um, Qwen models were better than, for example, um, um, um, um, Cl-Claude models, I think.
- 13:09
And sampling multiple answers, again, helped in this case. Synthetic rewriting and task augmentation, um, is something we thought will work, but it didn't very work-- uh, work very well.
- 13:21
And the other thing is, like, in, in this whole process of building this OpenThoughts agent, SFT still contributed a lot to the gains. Um, RL was kind of, you know, it's very compute-intensive and for, for the last few per-few percentages, it really helped.
- 13:38
Uh, but, but, you know, in many of the situations, for example, in enterprises, SFT actually works pretty well, right?
- 13:46
And here is one concrete example I wanted to share on, um, uh, actually deploying something to production, right, by post-training. So we have seen a lot of people talk about post-training, but in enterprise settings we haven't seen a lot of successes, at least I haven't seen, uh, that.
- 14:04
Here is a very concrete example of, uh, with Intuit, there is, uh, this app called Credit Karma, which if you install, there is a page wh-- uh, place where you can, uh...
- 14:13
The, the, the, the app gives you a reasoning as to why a credit card has been recommended, and this you can prompt a model to do this. But one of the reasons-- one of the places where it fails is that the, you know...
- 14:27
It, it-- it's not always compliant, so you have to have a long list of rules to make sure the responses are compliant, and that actually blows up the latency.
- 14:37
So answer here is, like, you want to curate data and post-train, right? Seems kind of straightforward, but one of the things that we ran into is, um, the dataset can be quite impa-- imbalanced and lo-lots and lots of places, for example, you will have zero percent APR, and the model after fi-fine-tuning can kind of hallucinate the, these,
- 14:59
uh, numbers. So this, again, kind of ties back to what Ross talked about some time back with respect to the tags, and we kind-- uh, we, we created this specific, uh, curation recipe where instead of just having these, uh, question-- the, the prompt response pairs in plain language, we added these, uh, tags, which helped the model to
- 15:23
focus on, you know, uh, the, the kind of form rather than the specific numbers itself, and that gave a big boost. And, uh, we, we saw that, um, the, the overall, the compliance metrics improved, the latency improved, the throughput improved, and eventually, you know, they, they are able to own the model, right?
- 15:42
As frontier models improve, they don't need to kind of go and, um, um, uh, update it. And also, as we see now, the, uh, [clears throat] frontier models are also getting more and more expensive and, you know, th-this kind of give, gives them a very good way for owning the model and also, um, lowering the costs.
- 16:04
I think with that, I want to briefly touch upon, um, uh, you know, Curator, the tooling that we had built last year, um, which is for curating reasoning data.
- 16:16
So, um, what it does is you can basically, uh, you know, um, specify the... You, you can either go with, say, a Hugging Face dataset where you have various prompts or, uh, in many situations you may have collected logs and you want to, uh, get the responses and fine-tune a model.
- 16:36
So this Curator kind of makes it pretty easy to do that. And it comes with the integration with, you know, uh, Tinker and Fireworks, and this, this is again, the tool that we used, um, originally for curating OpenThoughts.
- 16:52
And here is a very, very detailed diagram of what we are building today, but, uh, this again connects back to, um, what Ross was talking about, where he was talking about algorithms, uh, environments, and compute, right?
- 17:06
So it, it feels like, you know, we are kind of converging on something very similar. So if you think about, uh, the, the stack that is needed to, say, not just curate these RL environments, but to post-train models, one of the things you need is obviously handle on, like, how do you build these RL environments?
- 17:27
How do you measure the quality? How do you track the different versions and so on? So that's one of the layers. And below that you want various infrastructure to, uh, um, sandbo-- u-use sandboxes, right?
- 17:39
So spin up the rollouts-- to, to spin up the sandboxes to generate rollouts. And especially if you have long-horizon rollouts, then maybe at some point you need to do a checkpointing, and then you need to be able to snapshot or roll back to something else, right?
- 17:53
So that's the other, uh, the, the lower level, uh, you know, compute and orchestration. And at the top, I have been giving examples on post-training. So there is all this, uh, layer around, like how do you do SFT, how do you do RL, and so on.
- 18:07
But there is also this method called JEPA, which is around, uh, which is on prompt optimization. I don't know if you, if you guys have heard of it, but you can use LLMs itself to, uh, to, to kind of optimize the prompts based on reflection.
- 18:22
Um, so that also works pretty well for updating the system prompts and also the harnesses. So this is kind of, I feel like, you know, the, the new architecture or the new reference, uh, stack for how, how at least we are building and how many others are building, um, the, the stack on how to build the RLMs
- 18:42
and then also post-train agents. I think with that, uh, I'll, uh, end the talk and, uh, you know, happy to take questions offline. [audience applauding] [upbeat music]