AI Engineer World's Fair 2026

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

Read the talk

From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

Tarun Sunkaraneni walks through a multimodal training pipeline whose bottleneck moves from serial image loading to batch timing, tensor copies and finally a single node’s network card.

From a talk by Tarun Sunkaraneni

At a glance

Ideas worth remembering

  • A batch that needs more than 100 seconds of fetching and only 15–20 seconds of training is primarily a data-pipeline optimization problem.

  • Use concurrency that fits the work: asynchronous requests for I/O, threads where operations release the GIL, and processes for separate CPU execution.

  • Prefetching hides preparation behind training when the producer stays ahead, while adding queue startup and checkpoint-resumption complexity.

  • Passing object references lets training ranks retrieve large tensors directly, avoiding repeated copies through the coordinator.

  • Evaluate configuration at the intended scale: spread scheduling and zero-copy retrieval gave no small-scale benefit here, but together added 50% throughput at large scale.

Before the GPU does any math

An expensive GPU can spend most of its training run waiting for images. In Tarun Sunkaraneni’s experiment, the starting pipeline spends about 85% of its time acquiring data, leaving GPU utilization at roughly 15%. The first question is therefore practical: is the training loop actually blocked on GPU computation, or on the work needed to feed it?

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 210 secondsOpen full source frame
The slide highlights roughly 85% baseline data wait and less than 10% after optimization.

Multimodal training adds substantial work before a forward pass. Images need loading, resizing and normalization; audio needs loading and cropping; video combines these kinds of preparation. A cluster may have many more CPUs than GPUs, yet still leave its GPUs hungry if those CPUs and the input pipeline cannot deliver prepared examples fast enough.

The talk uses a wait-time ratio to track that problem:

Wait-time ratio = (preparation time + transfer time) / (preparation time + transfer time + training time).

The numerator represents time spent preparing and transferring data; the denominator includes training as well. Once stages overlap, the useful distinction is between preparation happening in the background and preparation that still makes the trainer wait. Sunkaraneni introduces an overall improvement from 15% to 90% utilization, with data waiting below 10%; the intermediate results later in the talk describe specific stages of that progression.

The workload is vision-language fine-tuning with multi-turn conversations. Its input path has three jobs before GPU computation begins: load the data, transform it, and transport it. Keeping those jobs distinct makes it possible to see which change removes which delay.

0:130:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:13 · section reference included

A batch of 256 waits for one image after another

The dataset uses JSONL records with relative image links, with the data stored in S3. One example points to an image described as a blue bus. Follow that example through the pipeline: its image must arrive from object storage, become a decoded and normalized image array, and reach the training process alongside the text.

A data generator handles loading, transformation and transport. A coordinator partitions examples among Megatron data-parallel ranks—the training processes that receive their respective portions of the data, collate them, and perform forward and backward passes. Megatron’s weight optimization is treated as a working black box. The changes in this experiment happen upstream of it.

In the baseline, a request for 256 examples triggers serial image fetching and serial processing. The bus image waits its turn among the other examples before the whole batch can reach the trainer. Training takes 15–20 seconds, while fetching the data takes more than 100 seconds. The profile is dominated by S3 loads, followed by image loading through PIL. That imbalance explains why optimizing GPU kernels would leave the largest delay intact.

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 352 secondsOpen full source frame
The baseline diagram depicts images fetched and processed one by one before training.
4:085:08
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:08 · section reference included

Match concurrency to the work

Each sample has two broad costs: fetching its images, then decoding them, applying vision preprocessing and tokenizing its text. Samples are independent, so the bus image does not need another example’s result before it can be prepared. A worker pool can work on many examples concurrently instead of carrying out the entire batch one sample at a time.

The choice of concurrency depends on what consumes the time:

  • Asynchronous I/O: Network requests spend time waiting for responses. Async programming lets other requests progress during those waits.
  • Threads: CPU-heavy operations can benefit when their implementation releases Python’s Global Interpreter Lock, or GIL. Whether a particular tokenizer, image transform or dataframe operation releases it needs investigation.
  • Processes: Separate processes can use multiple CPU cores and isolate workloads, at the cost of a heavier execution and communication setup.

The chosen combination is asyncio for fetching images and Ray actors for processing them. Async fetching overlaps remote waits; actors supply separate worker processes and APIs for communicating with them. This avoids building the worker communication machinery from scratch while giving image preparation multiple processes in which to run.

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 510 secondsOpen full source frame
The slide diagrams a worker based pipeline for concurrent data preparation.

The reported fetch time falls to roughly 40 seconds per batch, and Sunkaraneni reports a 25% speedup. Those figures describe the experiment’s different observations; the fetch-time reduction alone should not be read as an equivalent end-to-end speedup. More importantly, the trainer still asks for a batch before the workers begin obtaining it. Concurrency shortens the wait, but leaves the request on the critical path.

6:076:37
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:07 · section reference included

Prepare the next batch while the current one trains

Prefetching changes when preparation starts. A background thread fills a queue ahead of the consumer, rather than waiting for the trainer’s next request. The producer must keep ahead of the trainer’s consumption rate. For the bus example, fetching and transformation now happen before its batch is requested; when the trainer needs it, the prepared example is already available.

What delay does the queue remove? The flow below separates the producer’s early work from the trainer’s later request. The queue creates a place for completed preparation to wait, allowing data work and training to overlap. In the reported run, fetching and processing contribute close to zero seconds of waiting per batch because they have already happened—not because image preparation has become free.

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 570 secondsOpen full source frame
The slide shows a producer preparing batches ahead of the training consumer.

That timing change introduces two costs. The queue needs a cold-start period to fill, and checkpointing becomes more intricate because model state and the data loader’s progress must support a consistent resumption. Prefetching pays off after the producer gets ahead; it does not eliminate startup or the need to manage that extra progress.

How it fits togetherMove preparation ahead of the request

Fetch and process examples ahead of demand.

The producer fills a queue in advance; the trainer consumes prepared examples while preparation continues.

9:069:36
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:06 · section reference included

Stop sending the image array through the coordinator

Once fetching and transformation stop making the trainer wait, transport becomes conspicuous. Each prepared image array is megabytes per sample. Returning samples by value means pickling them on a worker, copying them to the coordinator, then serializing them again for the trainer. The bus image’s contents now take an expensive detour through a process that only needs to decide where the example belongs. Repeated copies can also push the data generator out of memory.

Ray’s object store changes what travels through the coordinator. Each sample carries an object reference; the driver holds that reference rather than the large tensor itself. Megatron’s data-parallel ranks retrieve the data from the store as needed, so the tensor no longer round-trips through the head node. This removes an unnecessary intermediate transfer while preserving the coordinator’s partitioning role.

Where do the large bytes go, and where does only a reference go? The diagram makes that distinction visible. The worker stores the prepared array; the coordinator handles its reference; the training rank retrieves the array directly. It avoids the coordinator detour, although retrieval still has a cost—a distinction that matters when the pipeline scales.

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 690 secondsOpen full source frame
The diagram shows a worker placing data in an object store while the coordinator passes a reference and training ranks retrieve it.

Using the store also removes an extraneous object-put operation introduced by the prefetcher. Communication costs fall, and the intermediate pipeline reaches a wait-time ratio of 20%, down from 85%. The reported utilization graph’s sawtooth pattern becomes much less pronounced: repeated periods of starvation have been reduced. Concurrency, early preparation and fewer copies each address a different part of that change.

How it fits togetherCoordinate references, retrieve tensors directly

Prepare the image array.

Large image arrays bypass the coordinator; object references carry the information needed to arrange their use.

10:3711:08
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:37 · section reference included

At production scale, locality overloads one network card

The same pipeline starts waiting more when scaled to a production run. The workload now has four times the data-parallel ranks, about 200 data streams, and four times the batch size. The generator and worker pool are colocated to make transfers fast, but concentrated processing also concentrates traffic. A node’s network interface card cannot keep up with the training pipeline’s demand.

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 780 secondsOpen full source frame
The slide labels network interface card saturation as the scaling bottleneck.

Investigating framework behavior reveals that Ray’s scheduling in this setup prefers colocating workers. That preference can exhaust one machine before making useful use of another. Spreading workers across nodes distributes image processing and storage; the data-parallel ranks then retrieve those prepared objects lazily, when needed. The tradeoff has changed: keeping work close together is less valuable once the shared network card becomes the limiting resource.

Two settings make the difference at this scale:

  • Spread scheduling: Place workers across multiple nodes so image processing and storage do not funnel through one overloaded machine.
  • Zero-copy retrieval: Reduce copying when retrieving prepared data, complementing the earlier change from returning arrays by value to passing object references.

Neither setting helps in the small-scale experiment, but together they produce an additional 50% throughput at large scale. Their value depends on which resource is saturated, so testing only the small workload would have led to the wrong configuration decision for production. Sunkaraneni’s advice is to ablate settings across the scaling ladder: compare their effects at multiple workload sizes, including combinations whose benefits may appear together.

12:5113:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:51 · section reference included

Each fix changes what needs profiling next

The completed pipeline prepares independent examples concurrently, gets ahead of requests, passes references instead of repeatedly copying tensors, and distributes workers when concentrated retrieval traffic becomes a problem. Its improvement comes from changing the path to the model while leaving Megatron’s training work as the black box established at the start.

The ending’s practical lesson is to keep asking what the run is waiting on: CPU work, I/O, GPU computation, memory or retrieval. Solving one bottleneck exposes another. Parallelism and prefetching are useful starting moves, but framework semantics—how results are returned, where workers are placed, how objects are retrieved—determine whether those moves keep paying off. Profile the baseline, then profile the changed system at the scale that matters.

Selected presentation frame from From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model at 960 secondsOpen full source frame
The slide lists questions for diagnosing whether a run is CPU, GPU, I/O or compute bound.
4:388:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:14 · section reference included

Read the complete timestamped transcript
  1. 0:13

    Hello, everyone. Thank you for joining me here. Um, yeah, today I'm gonna talk about Ray Actors, Vision Tokens, and the GIL. Um, engineering a multimodal SFT data pipeline that keeps the GPUs busy. So there's a lot of talk about, you know, how to optimize training paradigms. Um, but I wanna give a different perspective on the whole ecosystem of

  2. 0:43

    what it takes to get a model trained. Um, and so the question I'm gonna be asking here is, does your training loop really depend on the GPUs right now? Is it blocked on the GPUs? And I think we're gonna learn that-- We're gonna see in this experiment that the hardest part of optimizing multimodal train throughput is not necessarily kernel optimization. It is making sure that the D-- GPUs are constantly fed with data. So let's go over the overview of what this topic, what this talk will be

  3. 1:12

    about anyway. So we will establish a baseline multimodal training pipeline, um, and then we will encounter bottlenecks and then address them one at a time. So loosely, they are, one, concurrency. How fast can you process one image before feeding it to the model? Two, prefetching. How can we make sure that we have data on demand? Three, um, how can we minimize data copies across different processes? And four,

  4. 1:42

    fr-- uh, the additional framework optimizations, um, so just tweaking your bells and whistles. Um, yeah, so let's set up our premise. We all know that GPUs incur expensive costs and are generally want to maximize their utilization while we have them at hand. Multimodal processing is different than text-only, um, because there's a lot of CPU-intensive operations that have to happen before you actually

  5. 2:12

    feed the model, um, the data. So whether that is, um, you know, loading, resizing, or normalizing images using, like, a library like PIL. For audio, you'd do a lot of loading and cropping. You can use, like, NumPy arrays. And then videos are some sort of combination of both. Um, and then it's also true that usually your training clusters have a lot more CPUs than GPUs at hand, yet the GPU is what is starving in this pipeline.

  6. 2:42

    So today we're gonna introduce this equation called, like, wait time ratio, which is generally, I formulate it as, like, the time it takes to prep your data and transfer it to the GPUs. Um, and then d-we divide all of that by the time it takes to do both of those and the train as well. So basically, the denominator is, like, your total pipeline. Um, and we're gonna see that during our run, we will go from a utilization of fifteen percent to utilization of ninety

  7. 3:12

    percent. So our baseline has, like, eighty-five percent of its-- spends eighty-five percent of its time, um, backed on data acquisition, whereas after we do these optimizations, less than ten. And so for the scope of this talk, GPUs are not the goal. It's making sure that the data is never the bottleneck is the goal. And so here's a little bit of background as to, like, what our stack will look like. Um, we kinda take, like, a

  8. 3:42

    Qwen3-VL base model, um, kind of try to train it on the paradigms of ChatVQA and, like, multi-turn VLM, uh, conversations. So to recap, before the GPUs do any math, um, we have to load the data, we have to transform the data, and then we have to, we have to transport it to the GPUs.

  9. 4:08

    Okay, this is a, this is a overview of what our data actually looks like. So we will be using JSONL files, um, and then images are linked relatively, uh, to where the example lives. And then, uh, we're gonna assume that our data is just stored in S3. Um, in this example, for example, it seems like exa-- that image three three four seven one.jpeg seems to be a picture of, like, a blue bus. Um, yeah, so this is our system

  10. 4:38

    layout. On the left, we have a data generator, which is responsible for the data loading, data transforming, and tra-- data transporting. Um, and the-- that's, like, labeled as, like, fan-out to the Mega-Megatron data parallel ranks. Um, so the parallel ranks receive partitions of the data as decided by the coordinator. Um, it collates it, performs forward pass and backwards pass on this data. So Megatron takes care of optimizing the model weights across the DP ranks, and we're gonna assume

  11. 5:08

    that it's, like, a-- it's kinda-- it is kinda working as a black box. Um, so let's start with the naive baseline path. So for each example, we fetch the images from object store. Um, we kind of, like, we tokenize it, we resize it, rewrite-- resize the images, normalize them, and then feed them to the model. Um, yeah, and then when we do that, we can see that

  12. 5:38

    there's serial operations in the red, which is namely getting all the images one by one and then processing them one by one. So if your, if your trainer is asking for a batch size of, like, two fifty-six, then you have to serially process all these images. And then we can see that our train time is somewhere between 15 to 20 seconds. Uh, however, the data s- uh, the time spent to actually fetch the data is more than 100 seconds. So that's why we result in only 15%

  13. 6:07

    utilization of our GPUs. Um, and then you can even see a profiling flame graph that we spend most of the time doing S3 loads, and then you can see a long t- a little tail at the end about, like, loading the actual image on the PIL, uh, library. So okay, solution one, let's try to add some concurrency to this. Instead of those red singular p- uh, you know, serial processes, uh, let's try to parallelize it by kinda having, like, a worker pool. So when you request 256

  14. 6:37

    examples, maybe let's just process all of those at once. Um, yeah, so for example... for a sample, we have two costs. One is fetching the images, two is processing, so that's, like, decoding, vision pre-processing, and then tokenizing your text. Um, samples are very... are in-independent, so this just means it's, like, embarrassingly parallelizable. Um, and so yeah, you have a design decisions here. What forms of

  15. 7:07

    concurrency do you want? So there's asynchronous programming, uh, which is suitable for, like, I/O work, such as making SQL queries or, like, performing network requests. There's threads, uh, which you can use for, like, CPU-intensive, uh, workloads that relinquish the Python GIL, like, the Global Interpreter Lock. Um, so this... it's not straightforward what operations actually release a lock, so you have to do a little more of your investigation. But common uses are, like, using Hugging Face

  16. 7:37

    Transformers Tokenizers, using the PIL Transforms API, or, like, some Pandas operations. Processes, um, are more heavyweight, but also provide total isolation between your workloads. So you can leverage multiple cores on your machine as you finish your work. Um, yeah, doing a deep dive on the right solution would warrant its own talk, but, like, there's plenty of online resources that can help you decide for your use case. What's i-important for our solution is, is that we need to match our bottleneck,

  17. 8:07

    not simply going with a lever that's worked for us before. Um, so we'll choose asyncio for fetching images here, and then use Ray Actors as our processors, uh, as our process- multi-processing library. Um, Ray Actors are essentially the same as, like, general processes, but they offer us APIs to communicate across workers that just make it so that we don't have to reinvent the wheel. So yeah, your process is only as slow as your slowest link. Um, and we can see that our wait time ratio has gone

  18. 8:37

    down, um, on the second example, which is in the pink. Uh, you can see that the time spent f-fetching data per batch has gone basically, yeah, around, like, around, like, 40, 40 seconds. Um, so this is, this gives us a 25% speed up. However, you... yeah, that's, uh... as you can see, the... there's a huge gap that starts generating, especially as we're going through man-many steps. Um, so can we somehow limit our bottleneck

  19. 9:06

    of waiting on the trainer to request before provisioning our worker pool to fetch the data? So can we just, like, preemptively fetch our data? Um, so we came up with solution two, which was that we have a background thread that fills up a queue ahead of time, and it's always stays ahead of the consumer, which means that you... your producer is producing examples at a multiplier of your tra- of your consumer, and so the

  20. 9:36

    trainer is never bottlenecked on actually requesting the data. When it requests data, it's, it is available on demand. There's a caveat to prefetching. It makes your model and data loader checkpointing and resumption, like, a little intricate of a process, and it can also suffer from a cold start period in which your queue is actually getting filled up, uh, so it can be ready for training. Um, but a- if you work around those, those details, I think prefetching is a, is a great idea.

  21. 10:07

    Um, yeah, so I think making sure your producer is always ahead of your consumer will be... will pay great dividends for you here. Um, unsurprisingly, we achieved close to zero seconds of data fetching and processing per batch because we've already fetched data and processed it way before the trainer even needs it. Um, and so we've effectively solved both the data load and the transform bottlenecks. So after these two mitigations, we thought we were done. Throughput up, GPUs

  22. 10:37

    fed. However, we observed that the data transport costs still linger. In fact, they've gone up. So every sample transfers a large amount of, like, image arrays across process boundaries multiple times. Every sample is pickled on the worker, copied to the coordinator, and then re-pickled onto the trainer. Per sample, per step, we have a memcpy that nobody really asked for. So what ha- how this manifests is that your data generator can often OOM,

  23. 11:08

    um, and so multimodal tensors are huge. They're megabytes per sample. Um, and so by default, every sample is returned by value, which means this, that these expensive arrays are copied over from process to process. So a fix for this, you can come up with a object store. Um, this is a nice convenience method of... that Ray provides you. Each sample carries a object reference, which is just a pointer to the data, um, and then the driver just holds the pointer. The

  24. 11:38

    tensor never round trips to the head node. Um, and the Megatron data parallel ranks, they can just directly call the Ray store to fetch the data as they need it.

  25. 11:50

    So ev-with every choice of your parallelism and concurrency solution, you always have to think about the commu- added communication cost. So plan wisely around that. Okay. Yeah, and we were able to eliminate a extraneous, like, put object, uh, kinda call that our prefetcher solution had that we did not have to deal with by using the Ray store. And we can see that our, our communication costs go down, and the ratio of the time spent waiting also

  26. 12:20

    goes down. So to recap, we took a serial baseline, we added concurrency for data loading and processing, we added prefetching to stay ahead of training requests, and we added the Ray store to minimize data transfer. We w-went from a baseline wait ratio of 85% to 20%. And you can see that the sawtooth, uh, pattern that you see in the top graph is alleviated, uh, to a large extent, which means that our utilization of the GPUs has gone up.

  27. 12:51

    So when we're scaling up, now that we have, like, a good multimodal training pipeline, let's try to scale it up to, like, a production run. While doing that, we observed our wait time ratio for the same pipeline went up when all we did was scale some numbers. So we just added four times the data parallel ranks. We added, like, 200 data streams, and then batch size is X by four. Um, so it turns out that since the data generator and the worker pool are co-located for

  28. 13:21

    fast data transfer,

  29. 13:25

    what happens is that there's a bottleneck on the network interface card of that, of that node. Uh, which means that while we're packing all of this working and processing into specific nodes, that their network throughput is not able to keep up with the demand of the training pipeline.

  30. 13:45

    Um, so at this point, we did a little bit of digging into the frameworks that we are using, and we saw that by default, Ray schedules workers on the s-- it prefers to schedule workers on the same nodes and co-locate workers, which means that it's easy to bottleneck one machine before you move to another machine. So what we did was consider a possibility where we disaggregate this whole process where workers are spread across

  31. 14:15

    multiple nodes, um, so they can do their own image processing and storing, and in a lazy, just-in-time manner, are retrieved and used by our DP ranks. So yeah, to re-- And then so these are the two parameter, two f-flags that we kinda tuned. One is the zero-copy retrieval, and second is a spread scheduling mechanism. Um, neither of these actually helped in a small-scale experiment, but together they had, um, a

  32. 14:44

    synergistic effect in a large-scale experiment. Um, yeah, so s-config can lose... A config can do poorly at a small scale, but can do really good at a larger scale. So you should always think about how to ablate across your, your scaling ladders. And these were the two flags that ended up making a huge difference, uh, in how, how fast our data processing pipeline was. And so now this is the final pipeline that we've accomplished. A concurrent, eager,

  33. 15:14

    um, and a data pro-- uh, access pattern-aware system that shares the training across loads and enables that GPU stay warm while providing smoothness and stability in training. Overall, we achieved a 50% additional throughput at scale just from these two flags. It's the same knob, opposite verdict, set by the bottleneck. So these are the learnings we get from this talk, if nothing else. Um, GPUs are expensive, but sometimes the

  34. 15:44

    biggest battles happen outside of the GPU. So first ask yourself, are you CPU-bound or GPU-bound? Are you IO-bound or compute-bound? And then solving one bottleneck often exposes another. So first it was throughput for us, then it became wait time, then it became memory, then it became re-retrieval cost. Um, classic moves are not the end. You know, parallelism and prefetching matter, and I think a lot of content has, like, covered this. Uh, but so do the semantics of the

  35. 16:14

    frameworks that you're using and the, and the constraints that you're dealing with. And scale always changes the answer. Tuning parameters like zero-copy waste, um, and spreading workers become important only at larger scales, and they may not manifest as obvious improvements at smaller levels. S-overall, profiling your baseline helps you move the needle. Um, yeah, thank you everyone for attending the talk. Uh, if you have any questions, I'm happy to answer them.

  36. 16:44

    Thanks.