AI Engineer Code 2025
Lessons from Generating 12 Trillion Synthetic Tokens — Bogdan Gaza, DatologyAI
Read the talk
Lessons from Generating 12 Trillion Synthetic Tokens
Bogdan Gaza explains how DatologyAI seeds synthetic data with curated documents, then keeps trillion-token generation moving through metadata batching, recoverable jobs, coordinated scheduling and inference tuning.
From a talk by Bogdan Gaza
At a glance
Ideas worth remembering
BeyondWeb seeds generation with curated documents, then rephrases or restructures them to expand useful training material.
Metadata discovery can dominate startup at trillion-token scale. Batched S3 listing reduced DatologyAI’s estimated 9–11-day approach to about two hours of API processing.
Right-sized partitions and checkpoints limit repeated work after GPU failures; resumable execution and idempotent writes make retries practical.
Cross-cluster scheduling must coordinate a job’s CPU head and GPU workers, because separate pools can each have capacity without satisfying the whole job.
Inference parameter sweeps produced about 40% more throughput in the reported workload, while domain-specific data recipes remain an active experimental question.
Pretraining runs grow beyond the available data
Large pretraining runs need more useful data than the web can readily supply. That is the starting problem in Bogdan Gaza’s account of generating about 12 trillion synthetic tokens across web, math and code at DatologyAI. Synthetic generation supplements a finite source of training material, but producing that much text creates another problem: keeping expensive GPUs working throughout a massive data pipeline.
DatologyAI’s role begins before generation. Gaza, its co-founder and CTO, describes the service as an oil refinery for pretraining and mid-training datasets: customers bring substantial amounts of data, curation identifies the highest-quality points, and a synthetic data recipe expands those selected inputs. The selection step matters because the documents chosen here become the material the generator will transform.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
BeyondWeb expands selected documents through rephrasing
BeyondWeb is DatologyAI’s synthetic data recipe. The opening comparisons ask whether better training data can reduce the model size or token budget needed to reach a given accuracy. Models trained at 1 billion, 3 billion and 8 billion parameters provide the first comparison: the presented 3-billion-parameter BeyondWeb result matches the 8-billion-parameter Nemotron-Synth result.
The second comparison varies training tokens rather than model size. Gaza reports matching the accuracy of Nemotron-Synth with about 2.7 times fewer tokens and Cosmopedia with about 5.3 times fewer tokens. These are results from the presented comparisons, rather than universal reductions in training cost; the recording does not specify the complete evaluation conditions. Gaza describes the results as dating from around late summer 2025 and says the recipe has continued to improve.
The recipe starts with an existing high-quality document. Asking a model to invent a synthetic data point without a seed tends, in Gaza’s explanation, to draw from the common modes of its original training distribution. Supplying a selected document gives generation a specific starting point. Prompts prefixed to that document request rephrasings, restructurings or reframings, and models running on GPUs produce the new material. The described inputs can be text or image-text data.
This makes curation and synthesis one connected process: identify useful source material, then generate alternate presentations of it. The GPU work is straightforward to describe at this level. Predicting how quickly the entire dataset will finish is harder, because actual throughput fluctuates and falls short of the ideal rate.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Bring research and product jobs into the same workflow
The original setup split synthetic generation from the product pipeline. Research ran generation, training and evaluations on an H100 cluster managed by Slurm; the core product pipeline ran on Kubernetes. Separate code paths slowed experiments, introduced errors and left GPUs underused. Generated data also required manual backfilling into the product’s asset tracker, its data catalog.
The move to Kubernetes retained separate clusters with different jobs. A control cluster holds orchestration and the asset tracker. A compute cluster runs Spark and Ray work using on-demand CPUs and, when appropriate, GPUs that do not require H100s. The H100 cluster now runs under EKS on HyperPod. An internally built scheduler coordinates work across these clusters, with Ray, KubeRay and vLLM forming the core synthetic generation stack.
Two infrastructure choices support deployment beyond this particular AWS setup. The pipeline runs on Kubernetes because customer environments may be AWS, GCP or on-premises. Data lives in S3 or storage that speaks S3-compatible APIs. The architecture therefore keeps a common orchestration approach while allowing compute to live in different environments.
What becomes possible when those jobs share one workflow? Curation can run Spark jobs, feed synthetic generation, launch training and then run evaluation. The flow below shows the experiment’s dependencies. Each stage supplies the next, allowing research and engineering to use the same infrastructure and code path rather than manually reconnecting outputs between systems.
Spark jobs select and prepare training data.
A shared scheduler connects stages that previously crossed separate research and product systems.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Metadata can leave the GPUs waiting for days
The first bottleneck appears before inference. Gaza estimates that the web supplies about 30 trillion usable text tokens, give or take, and names FineWeb, DCLM and Nemotron as starting datasets. At this scale, data stored in formats such as Parquet spans millions of partitions. Ray must discover and manage the corresponding storage metadata before workers can efficiently process the data.
Fetching metadata through many individual S3 requests makes startup depend on request overhead and rate limits. Gaza uses a roughly 10-millisecond request as an illustration, then describes a back-of-the-envelope estimate of 9–11 days to populate the metadata for the intended scale. That estimate is the warning: a generation fleet can have ample GPU capacity and still sit idle while the Ray head assembles the information needed to dispatch work.
The change was to batch metadata discovery through S3’s listing APIs, increasing the page size to about 1,000 entries per list. More metadata arrives per API call, reducing the repeated request work. Gaza reports about two hours of API processing time with this approach. The practical lesson is to plan how bucket state enters the execution system: accelerating inference does little for a job that cannot finish discovering its inputs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make an eight-hour partition survive a late failure
Once generation starts, hardware failure becomes part of normal operation. A machine can go down or a GPU can fail, leaving partial outputs and consuming cluster time without completing the assigned work. In this pipeline, Spark stitches prompts together with the documents to rephrase. Ray then runs GPU generation, with jobs ordinarily organized around a one-to-one mapping between input and output partitions.
Consider Gaza’s concrete failure example: one node takes eight hours to rephrase one file, then goes down in the final five minutes. Without saved intermediate progress, almost eight hours of work must be repeated. Retrying the job restores execution, but it does not recover the compute already spent. The size of the partition determines how much unfinished work is exposed to that failure.
Two controls reduce the loss:
- Right-sized partitions. Smaller units of work limit how much progress a failed partition can discard.
- Periodic checkpoints. Flushing intermediate progress to S3 preserves work before the entire partition finishes.
DatologyAI uses both. A failed job can retry with a tolerable amount of lost work instead of turning into a multi-day failure.
What changes at the moment of failure? The comparison below follows the same eight-hour job through its retry. Checkpointing preserves a completed prefix of the work; resumption must redo only work beyond the last saved point. Gaza also calls for idempotent checkpoint writes, describing a checkpoint file that can be overwritten, so retrying a save does not require a new logical output each time. The recording gives no checkpoint interval, so the remaining loss depends on that choice.
One partition requires eight hours.
The same failure near the end of an eight-hour partition causes different amounts of repeated work depending on whether progress has reached S3.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Schedule the CPU head and GPU workers together
Synthetic generation does not necessarily need fast networking between GPU nodes, which makes geographically distributed capacity useful. Curation can remain centralized while customers’ generation jobs run across several GPU locations. The difficulty shifts to coordinating the resources each job needs across those clusters.
Separate resource availability can produce a scheduling mismatch. One job may obtain its CPU resources while its GPU workers cannot be scheduled. A later job may find GPUs available but no CPU capacity. Either resource pool can look usable by itself while the combination needed for useful execution is unavailable. Gaza describes this as a bin-packing problem.
The solution keeps dedicated pools for different demands, including Spark CPUs and Ray heads, then coordinates the CPU and GPU requirements of a Ray job in a more atomic fashion. Separating pools makes the resource roles clearer; scheduling the head and workers together addresses the mismatch between independently allocated pieces. The talk describes this scheduling principle rather than a specific cross-cluster reservation algorithm.
The important unit for scheduling is therefore the working job: a CPU head and the GPU workers it needs. That requirement persists when capacity spans multiple providers or clouds. Increasing either pool alone will not resolve every queueing problem if the scheduler cannot assemble both sides at the same time.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Benchmark inference settings for the batch workload
The final bottleneck sits inside inference. DatologyAI runs substantial batch generation through vLLM and is also testing SGLang. This workload differs from online serving: the task is to process large volumes of generation work, so familiar serving settings need to be benchmarked against that workload.
Gaza recommends a dedicated benchmark harness and grid searches over inference parameters. Two named controls are batch size and speculative decoding configuration. The harness makes their throughput effects observable, allowing the team to choose settings from measured combinations rather than accepting defaults.
The reported gain is about 40% more throughput from flag tuning. That result belongs to the workload and configurations tested; the recording does not give the model, selected flags or baseline settings needed to reproduce it. The useful decision is to treat inference configuration as an experimental variable alongside the data recipe.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From 30 billion tokens to 12 trillion—and the next question
At the end of 2024, DatologyAI could produce around 30 billion synthetic tokens in limited Slurm runs. Gaza’s closing comparison describes later customer jobs producing about 7 trillion web tokens plus another 5 trillion math and code tokens. Metadata discovery, recoverable execution, coordinated CPU/GPU scheduling and inference sweeps together supported the move to this much larger production scale.
The ending leaves a substantive question open: how should the synthetic recipe change across domains? Web, multilingual, math, code and legal data may require different treatment, and Gaza identifies domain coverage as important without developing those differences here. The team continues to run ablations and search for better recipes. A pipeline that can curate, generate, train and evaluate in one workflow makes those experiments easier to carry through, while the volume generated alone does not settle which recipe works best for each domain.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
We're doing this. Hey, everyone. Welcome to my talk. I'm Bogdan. I'm one of the co-founders and the CTO of Datology. Today, we're gonna talk about engineering lessons from scaling synthetic data for trillion tokens scale pre-training. Um, let's see. So why do we need synthetic data in the first place? Well, turns out that with, uh, the scaling laws, we kinda hit a data wall. Um, you need to put exponentially more compute and exponentially more data to get, um, models that are getting kinda linearly better and better. So,
- 0:42
um, we're, we're at this point where pre-training runs are massive, so in order to do the, the kind of-- the ones that actually get to, I don't know, the mythos-level, um, capabilities, you need a lot of data, and a lot of it, uh, cannot be found on the web. So, um, internet has limited amount of data. Um, we can supplement it with high-quality synthetic data. So what are we gonna talk about today? We're gonna talk about, um, you know, how do we do this in the first place. We're gonna talk about-- a bit about
- 1:12
BeyondWeb, which is kind of our synthetic data recipe. Um, we're gonna talk about the engineering lessons from our prior experience here at Datology of running trillion tail-- trillion, um, token scale synthetic data runs. Um, we recently completed one for about twelve trillion tokens. That's across web, math, and code, and we're gonna talk about what we learned from running absolutely massive synthetic data runs. Um, and then we're gonna talk about kind of what's the-- w-why does this matter in the first place? Um,
- 1:42
awesome. So let's get started. Um, so the way that we solve synthetic data at Datology, and maybe let's start with a, with a quick intro about Datology. We're a kinda data curation service. Uh, think of us as the oil refinery for your kind of pre-training and mid-training datasets. We work with customers, both startups and enterprises, that come to us to, um, kinda help them with their pre-training and mid-training data needs. So in general, they have a lot of data. Um, they wanna get to the best datasets possible, and we help them identify the, the highest
- 2:12
quality data points and, um, once we do that, we pass them through our synthetic data recipe, the ones that we're gonna talk about today, um, to get absolutely massive datasets. So, um, that's a bit about us. Uh, in terms of BeyondWeb is, is our synthetic data recipe. Um, there's kind of many ways to do synthetic data but, um, kinda the, the, the way that we approach it is, um, quite different. So here's a few results from, from, um, kinda how we do this. So, uh, there's two plots. Maybe let's start with the first one. Um, on the X-axis, we have the model
- 2:42
size in billions of parameters. Um, one billion, three billion, eight billion, and then on the Y-axis we have, um, accuracy. So, um, we've taken kind of a, a number of datasets like Nemotron-Synth, Cosmopedia, QA WRAP, um, RedPajama, and then we've trained models at kinda one B, three B, and eight B scale together with, um, kinda the, the recipe we've put together for, for, uh, synthetic data which we call BeyondWeb. Um, and what we see is that, um, kinda our synthetic data recipe, um, kinda
- 3:12
matches and, and sometimes even, um, kind of, um... In this case, the, the three B, uh, data point matches the eight B data point, um, for Nemotron-Synth. So that means that you can train slightly smaller models for the same performance as you would if you train larger ones, uh, by using synthetic data. So, um, that's a bit about kind of the, the, the, the quality of, of our recipe. Um, same thing here when you, when you talk about performance. So on the X-axis we have, um, kinda the number of tokens in
- 3:42
billions, and in the Y-axis we have, uh, the, the average accuracy. Um, and what we show is that we match the, the, the performance of, um, kinda the, the, um, Nemotron-Synth and the Cosmopedia ones in about two point seven x and five point three x respectively, uh, less tokens. So, uh, if you wanna train a model to the same performance, you can do that, um, by using less data and, and overall less compute. Um, so at, at a high level what, what BeyondWeb is, is, is kinda how we think
- 4:12
about our synthetic data recipe. What we do is we do rephrasings, um, so where we, we, we kinda seed the rephra-- the, the synthetic data process by, um, kind of finding the highest quality data points in a customer's dataset, and then we rephrase or restructure these data points in, in, in various ways. Um, this is-- results are actually a, a bit old. I think the, the-- originally we published this in, uh, in, um, I think late summer '25, and at this point we're, we're much, much better than this. Um, and, and we're continuing to improve quite a bit. So wh- how do
- 4:42
you get synthetic data, right? In the, in the simplest form, we have some sort of target dataset. Again, we're doing rephrasings, re-- um, um, or reframings of existing documents. So the idea behind it is that, um, if you just prompt a model and say, "Hey, give me a, a synthetic data, um, data point," you, you're gonna just learn the, the modes of the distribution the data was, was origin trained on. So, uh, the, the way that we do it, as, as mentioned previously, is that we seed it with, um, some sort of existing document. So we identify high-quality documents in the customer's,
- 5:12
um, training datasets, and then we, we pass them through, through our pipeline. So at the simplest set, we have a, a target dataset. Imagine this can be text, this can be image text. Um, in, in the simplest forms it is just a, a bunch of prompts that, that we prefix to, um, a, a number of documents that we've identified. You have a bunch of GPUs, and then you, um, kind of use various models to generate synthetic data. Um, so in practice you have some sort of, um, expectations around how much time this is gonna take. But
- 5:42
in practice, this is a lot more wavy and, and it, it really-- it's really hard to match the, the kinda ideal throughput that you're trying to get of your, um, out of your GPUs. So how do we manage that? Well, well, where did we start? We started where, um, synthetic data wasn't integrated in our product. It was something that we would do in a research setting, uh, on a cluster that, that runs Slurm that is away from our core Kubernetes-based pipeline. Uh, we had this two-track system where kinda the product and the research team would, would not use the same type of code base. This was slow,
- 6:11
error-prone, um, limited in velocity, and kinda led to, uh, underutilizing our GPUs. We have a set number of GPUs like everybody out there, and in general we wanna be able to, to maximize the, the, the use of our GPUs by running this in a, in a cohesive way, both with training and synthetic data runs. Um- Um, so where do we start? Now, in general, uh, um, we, we like to run everything on, on Kubernetes. Um, a lot of our pipeline is either Ray or Spark. Today we're gonna talk about, um, kinda Ray and KubeRay and vLLM. Um, those
- 6:41
are, are kinda the core components of our synthetic data library, um, or synthetic data, um, kind of pipelines. So today what we do is we have an orchestration layer in one Kubernetes cluster. We, um, run Spark in another Kubernetes clusters, and, um, we, we keep the asset tracker, which is our, you know, data catalog in the original control cluster where everything, um, kinda gets coordinated and, um, into. Um, all the data is stored in S3, that's our storage layer. Um, as long as it s- speaks kinda S3-compatible APIs, we can run it. So in,
- 7:11
in some ways, we... everything needs to run in Kubernetes because our pipelines get deployed in all sorts of environments. Um, it can be AWS, it can be GCP, it can be on-prem, et cetera. We wanna make sure that, that, um, kinda we can run it in any, um, Kubernetes environment. And then separately we had this H100 cluster. Um, it ran synthetic, it ran training, it ran evals, kind of a number of things, all of them originally orchestrated by Slurm. Um, moving data in between these two clusters was manual, so you- whenever you generated synthetic data in Slurm, you had to backfill it into the, the asset tracker into
- 7:41
our data catalog. So where are we now? So, um, kinda same setup in some ways, where we have the, the original cluster, um, that, that does the control. We have the compute cluster that, that we run Spark jobs, and now we have this other H100 compute cluster. In, in our case, we run everything on AWS. We run it on HyperPod. HyperPod has this product called, um, EKS on HyperPod. So it's basically a bunch of H100 clusters on the same spine, all of them, um, orchestrated by, uh, by Kubernetes. So what we do is we continue to run Spark and, um, Ray jobs
- 8:11
on-demand capacity, on-demand CPUs, and sometimes on-demand GPUs, um, especially for, for things that we don't require H100s, we can do that. Um, and then we orchestrate it in a similar way that, that we've done before, where kinda Spark runs within its Kubernetes cluster, and then Ray runs within, um, the same cluster for, um, GPUs that are not H100s. And anything that we need to do for, um, H100 nodes, we go to the o- this other Kubernetes cluster, and we, we kinda, um, have a job scheduler that we built internally that, that orchestrates all of these things. Um,
- 8:42
if we do this, then this allows us to, to run fast experiments. Imagine that you can run, kinda, um, curate where you run a bunch of Spark jobs, you run a bunch of synthetic data, then you go and, and launch a training job, and then you run, uh, an eval job, and you can orchestrate all of this together in one workflow, which, um, for our research team is, is quite useful. And this is... puts the, the, the research flow and the engineering flow in, in the same lines, a- in the same infrastructure, in the same code, in the same path. Okay. So, um, we... W- where are we now? Right? Like, where... Oh, look, I have
- 9:12
a, a one-on-one with someone on the spine. Um, w- where, where we're at is, um, kind of synthetic data is, is a core thing to, to our, uh, workflows. Um, we ha- we run experiments all the time. We, we do this cross-cluster scheduling. It's a lot more faster, it's a lot more reliable, and we can run it at scale. So, um, we've recently done a massive, um, synthetic data run. Uh, let's talk a bit about, uh, the engineering lessons that, that we learned from this, right? So, um, w- what are the bottlenecks at, at trillion token scale?
- 9:42
Um, well, first and foremost, um, i- i- internet scale, um, for, for text data set is, is quite limited, right? Like, on, on, on the web, we can find about 30 trillion tokens, give or take. For, um, images, there is about two... 12.8 billion. In, in things like DataComp for, or, uh, for text, we usually start from existing high-quality data sets like FineWeb, DCLM, and Nemotron. Um, we can, we can, um, kinda put them all together, and we get this, these large data sets. Um, but at this scale, um,
- 10:12
u- usually everything gets stored in Parquet or some sort of, of, uh, o- of format. You have millions of data partitions. Um, if you're trying to fetch all of them from S3 at a high rate, you might get throttled. And then if you, especially if you run this on Ray, um, your, your Ray metadata is, becomes... and how you manage that becomes a bottleneck, and that leads to the, the worst thing possible, which is that you get, uh, idle GPUs. So bottleneck, um, that, that we initially seen is that we have, um, S3. We try to fetch things from our Ray head. Um, this
- 10:42
has some sort of metadata store. Um, this is kind of quite slow, and, and if you do one request that takes 10 milliseconds and, and, um, you're trying to, to do, I don't know, um, millions and millions of them at, I don't know, let's say, um, uh, if you need to do, um, kind of the metadata fetches originally, those are more expensive. Um, and there's also a rate limit. So if you need to do millions of them, you're gonna get rate limited, then your workers are gonna be quite slow. Um, so we kinda ran the, the, the back of the en- of the envelope math just to populate
- 11:12
the, the metadata for, um, trillions and trillions of tokens, it would take about nine to 11 days, uh, which is unacceptable, right? Like, you, you, if you're running this, you're probably running on millions of dollars in compute. You cannot have them, um, just waiting for you to fetch metadata for, for 11 days. So the, the, um, way that we've went past this is that, um, instead of, of kind of, um, kinda going to S3, you can actually do, um, requests that are, um, batched together. So instead of doing one at a time, you can batch them all
- 11:42
together. S3 supports this through their APIs, and then this way, um, you, you can, you can do this, uh, at, at a massive scale by, um, increasing the, the, um, the page size of your, of your batches to, to kind of the limit that, that S3 gives you, um, which is about a, a, a, a thousand, um, kind of requests in one list. And that gets you to about, um, kinda two hours of API, uh, processing time to get your, your metadata. So first lesson, um, make sure that you know how you manage your metadata and make sure
- 12:12
that you have, um, kind of thought through just getting the state of your S3 buckets into whatever synthetic data system that you have in order for you to, um, kind of be able to, to even start doing this at scale. So 11 days to two hours, uh, that's, that's the first one. Um, awesome. So the second one is GPU instability. So whenever you're running, um, these types of large scale synthetic data runs, of course you're gonna hit, um, some sort of bad GPUs, right?
- 12:42
So that means that the machine might go down, a GPU might be bad, something around, uh, those lines. Everything that you can expect when you try to run it at scale would happen. So, uh, that leads to lost progress, that leads to kinda wasted compute, partial outputs, um, if you, if you only pow- parse through half of the file. And of course, a- at the end of the day, you're gonna, uh, need more cluster time, which is always hard to get. Um, so the, the, the way that we do this in, in our pipeline is that we have a Spark job that outputs the, the, the prompts together
- 13:12
with the, the documents that we need to rephrase, that s- it, it stitches them together. Um, and then we start a Ray job that, like, and it actually uses the GPUs, right? Um, so the, the, the problem is that there's some sort of one-to-one mapping between input and output. That, that's, that's kinda how we, we usually, uh, structure these types of jobs. Um, in practice, if your partitions are, um, quite large, what happens is that, um, you're gonna do a bunch of, of, of, um, kind
- 13:42
of progress, and then if you- if, if a, a rephrasing job on one node for one file for one partition takes eight hours and the box goes down in the last five minutes of those eight hours, you're gonna lose a lot of time. So you need to think about, um, kinda... You either have to checkpoint, and whenever you do progress within the synthetic data, um, jobs, you need to periodically say, "Okay, let me stop," and, like, flush this back into S3. Or, um, you need to have partitions that are small
- 14:12
enough that, uh, you can tolerate that downtime. So, um, you need to, to think about either your partition size or, in general, kind of have some sort of checkpointing. Um, what we've done is, is a bit of both, where w- we've managed to rightsize our partition sizes and, uh, checkpoint as well when it comes to these things. So, um, basically, whenever, um, a job fails, it gets retried, but the amount of, um, kinda time that we lose is, is minimal, and it makes sense for the setup that we have. Um, so from multi-day job failures,
- 14:42
now we have recoverable retries, and it, it, it works, it works as, as expected. So just to recap, resumable execution, some sort of idempotency where whenever you're checkpointing can override the file that you're checkpoint, checkpointing into, and then you need to have some sort of mechanism for, for seamless recovery. Um, cross infrastructure orchestration. So in, in our case, the, the, the cluster where we originally run, run curation, in general, our pipeline runs in a centralized location, and then customers might have synthetic data, um, jobs that run in
- 15:12
various GPU locations throughout the world, right? You might... I- it is gonna be hard to get, especially for synthetic data, where you don't need necessarily fast networking in between your GPU nodes to get a lot of capacity in one place. So what happens then is that you need to be able to orchestrate all of these jobs, um, in, in various clusters. So it's not just one GPU cluster that you're working with, you're working with, um, kind of a number of them. Um, the problem here is that whenever you have, um... and you're trying to, to, to schedule, um, these jobs, is that you might
- 15:42
be able to, um, kinda schedule your CPU requirements for your Ray jobs, but not necessarily your, um, kinda GPU requirements, uh, in the GPU pool. So you need to think about kinda being able to schedule in some sort o- of fashion across these two things. Now, um, once you have the, the, the next job coming in, um, you have GPU pool capacity, but you don't have CPU pool capacity, right? So in general, whenever you're trying to do this, this type of cross infra, uh, or orchestration, you need to think about scheduling, and you need
- 16:12
to think about, um, kinda being able to schedule this effectively. Um, and in this case, as I mentioned, you, you, you will get to a point where you cannot schedule effectively. So what happens then is that, um, you, you, you kinda need to think about the, the, the ways in which you do this in a, in a proper fashion. So some sort of bin packing problem at the, at the end of the day. Uh, so how did we solve this? Um, so same setup as before. Um, we have this orchestrator that runs in a different cluster. We have the compute cluster that runs Spark and the,
- 16:42
the Ray jobs that mostly need CPU, and then we have the, the, the one that has H100s. Um, so whenever we schedule, um, we have dedicated pools for the specific resources that you need in those clusters. So if you have a, a, a Spark requirements, then you Spark... You might have a, a Spark pool that, that has CPUs, and then you, uh, dedicate special pools for, um, kinda the, the Ray head. So when you think about, um, some sort of a- atomic scheduling, you need to think about both your CPU and your GPU requirements in order to do this
- 17:12
optimally. So in this case, whenever we have more CPU, um, requirements that we need to schedule, because we've decoupled these two things, um, and, and then we can schedule the, the, um, head and the worker for Ray, um, uh, in a, in a more atomic fashion, then we don't have the problems as, as before. Um, and then everything scales as expected. So cross infra orchestration, um, you need to think about, uh, multiple providers, multiple, um, clouds that you run this in. You need to think about scheduling
- 17:41
and, um, make sure that whenever you schedule CPU and GPU resources, especially for the same type of job, you do that in a way that, that makes sense. Um, last up is inference optimizations. Um, so whenever you run vLLM or whenever you run SGLang, um, you need to think about the, the hyperparameters that you use for your jobs, right? And, um, what, what we found is that you need to be, um, quite disciplined about this. You need to be able to, to benchmark this and optimize and have a harness that just focuses on the vLLM parameters. You can get, um,
- 18:12
a lot of gains, um, in the case... in, in this case, uh, about 40% throughput gains just by tweaking the flags for vLLM, um, and SGLang, right? So in our case, we, we... when we're in vLLM, uh, right now we're actually testing SGLang as well. Uh, we do a lot of batch inference, um, which is very much different from, from the, the normal online one that everybody's used to. Um, but whenever you, you kind of orchestrate this through vLLM, you need to make sure that you have the, the right batch size, you have the right speculative decoding, um, set up,
- 18:42
et cetera. Um, if you do, uh, kind of a bunch of grid searches over the parameters that you need to optimize, turns out you can get quite a lot of, uh, of, of results in, in terms of throughput. So to recap- Um, whenever you're, you're, you're trying to do these things, make sure that, um, you think about kinda, uh, the, the scale that you're running at. You have a good sy- system for auto recovery whenever jobs go down. You need to orchestrate in some sort of atomic fashion across CPU and GPU resources. Um, and you need to make sure that, that you kind of sweep over your vLLM parameters. So what, what
- 19:12
this setup helped us achieve is, is quite straightforward. Um, and, and at the end of 2014... Sorry, 2024, uh, we can only output a- around 30 billion synthetic tokens in kinda limited runs within our SLURM cluster. Uh, our requirements for synthetic data went up by, by a bunch, and then, um, kind of end of last year, we were able to, to run this massive, uh, synthetic data jobs that, that, um, generate, in this case for web, about seven trillion tokens, and then for math and code, another five for a bunch of our customers.
- 19:42
Uh, we're continuing to ablate and, and kinda search for better synthetic data recipes. Um, we're actively working towards this, so, uh, keep an eye out. We'll be publishing more. And, um, I think domain coverage is quite important. I think, uh, something we haven't talked today is how does this differ in between web, multilingual, math, code, uh, legal data, et cetera. By the way, we're hiring, um, throughout the board for data infrastructure, cloud infrastructure, normal infrastructure, um, program managers, et cetera. So hit me up afterwards. Uh, thank you for the team that worked on this. And, uh, thanks everybody for attending my talk.