← All AI Engineer talks

AI Engineer World's Fair 2025

Maximize GPU Efficiency with Continuous Profiling for GPUs

About this talk

Matthias Loibl of Polar Signals explains continuous GPU profiling that combines NVIDIA NVML utilization, memory, power, temperature, clock-speed, and PCIe metrics with low-overhead Linux eBPF CPU profiles. He demonstrates correlating GPU underutilization with Python and CUDA call stacks in flame charts, measuring CUDA kernel execution time, and deploying the profiler through systemd, Docker, or Kubernetes.

Chapters

  1. 0:00Continuous profiling fundamentals and sampling overhead
  2. 2:53eBPF, speaker introduction, and GPU profiling preview
  3. 4:43NVML utilization, power, temperature, and PCIe metrics
  4. 6:24Correlating GPU metrics with CPU profiles and flame charts
  5. 8:13GPU time profiling and CUDA kernel duration
  6. 10:09Deployment with systemd, Docker, and Kubernetes

Talk transcript

  1. 0:00

    [upbeat music] Great. Uh, thank you for, for coming.

  2. 0:17

    Um, I'm gonna talk about maximizing GPU efficiency with continuous profiling for GPUs. Um, so what is profiling? Profiling is pretty much as old as programming. I think it was, like, firstly-- uh, first done in, like, the 1970s.

  3. 0:32

    I think some IBM folks were, were trying to figure out what was happening on their computers back then. So it's been around basically forever in computer science. And, uh, what are we doing with profiling?

  4. 0:46

    We are profiling, um, basically, uh, anything that, that we can to, [chuckles] to inform our, um, view of the world. We wanna see, like, the memory or CPU and GPU, um, time spent.

  5. 0:59

    We want to see the usage of the individual instructions and the frequency and duration of these function calls. So, um, yeah, a lot of different approaches to profiling, but, um, yeah, it's generally speaking, super, super important to, to performance engineering.

  6. 1:16

    So why would we do this? Uh, obviously to improve performance, and also we can save money. So if we, um, improve our, um, software by, like, ten percent, we might be able to just, like, turn off ten percent of our, like, servers and save a bunch of money, right?

  7. 1:32

    Um, so that would be great. And, um, there are two different kinds of, of profiling typically, um, that, that we're seeing these days. So one is tracing profiling, um, and that is you record each and every e-event all the time constantly.

  8. 1:48

    But, um, obviously, that's great for, like, getting, like, the best, uh, possible, uh, view onto the system, but it's, like, pretty high cost, um, and generates a lot of data, so it's, like, hard to, to do, uh, continuously.

  9. 2:05

    And that is why we're, uh, doing sampled, uh, profiling. So what we do is basically we sample for a certain duration, like ten seconds, and we, uh, only sample a hundred times per second or, like, twenty times per second, et cetera.

  10. 2:19

    Uh, you can tweak that, uh, how often do you wanna profile, um, and, and, and sample. So, like, a hundred times per second isn't that much for a CPU, and that's why you get, like, less than, like a percent overhead on, on the CPU and, like, only like four megabytes of overhead, uh, for the memory profiling.

  11. 2:38

    Um, you will most definitely miss things, but if you do it always on, um, you will eventually see most of the relevant things, right? Like, one stack that executed once isn't, like, relevant to us anyway.

  12. 2:53

    Like, we wanna see the, the big picture.

  13. 2:57

    Um, so yeah, this is basically what we're, what, what we're doing. We, we see, like, the stacks on the left-hand side executing, and these are, like, the functions that are calling each other, and we are just, um, like twenty times per second or a hundred times per second, taking, taking note, um, on what exact, uh, stack we're

  14. 3:16

    seeing on the CPU, um, or which stack is, like, allocating, uh, et cetera.

  15. 3:24

    Yeah, and that allows us to, like, do it always on, do it in production. Um, your machine is not the production environment, so it is pretty important to be able to do this in production and actually see what's happening, uh, out there in the real world and do it with lo-- uh, low overhead.

  16. 3:42

    And we are actually using, uh, Linux eBPF.

  17. 3:46

    And because we're u-using something, um, that the kernel is doing, we, we don't even have to, uh, change any of your, uh, applications. That means, um, you start one thing and it will start, um, profiling all of your applications.

  18. 4:02

    So you don't really have to instrument. Quickly about me. I'm Matthias Loibl, flew in from, uh, Berlin, Germany, and I'm the director of Polar Signals Cloud, and I'm also a maintainer of Prometheus, Prometheus Operator, Parca as the open source, uh, version of all of what I'm talking about today, and some other projects.

  19. 4:26

    So, um, we are basically here for, like, GPUs, right? And we just earlier this year, uh, after, like, working on CPU and memory profiling for the last three or four years, um, started, um, a preview on GPU profiling.

  20. 4:43

    So I'm gonna talk about this today and, uh, why, why we think it's pretty, pretty great. Um, as you can see in this, uh, screenshot, um, we're talking to NVIDIA NVML to get these metrics out of, uh, your GPU.

  21. 4:58

    So we can see in the blue ch-- the blue line on top, we can see the overall utilization of the node, and then the orange, uh, line is one particular process on the GPU.

  22. 5:11

    Um, so we can see-- Over here, we can see the process ID. So we see individual processes, but we also see the overall, uh, nodes utilization, um, further down the memory utilization and the clock speed, et cetera.

  23. 5:26

    And that will kind of inform, um, where we wanna look at, um, the performance of our system, right? So sometimes we can see the utilization drop down, and that might be something that we wanna investigate to really make sure that we are using, uh, our GPUs to the fullest.

  24. 5:42

    Uh, just couple of more metrics we are collecting. So there's, like, the power utilization, and the dashed line is the power limit, and then the temperature. Temperature sometimes is important because, like, eventually, if you're, like, always at, like, eighty degrees Celsius, you're, you're gonna get throttled, um, by the GPU, um, quite significantly.

  25. 6:03

    And then obviously PCIe throughput Um, it's interesting, are you bound by the data you are transferring between CPU and GPU? Perfect. Yeah, so, uh, just to repeat, like, um, the negative one is, uh, receiving whereas the, um, positive ones are sending 10 megabytes per second, uh, through PCIe.

  26. 6:24

    And then we can, uh, use all of those metrics to correlate from the CPU prof- uh, from the s- uh, GPU metrics to the CPU, um, profiles that we're storing.

  27. 6:34

    So we're, like, collecting, like we have done the last f- three or four years, uh, using eBPF, those, uh, CPU stacks. Um, and we wanna, like, see what is happening on the CPU.

  28. 6:46

    So in this case we might wanna look at a particular stack on, uh, CPU zero, um, right before the end because there was some activity, for example. So we can drag and drop and select a particular, uh, time range, and then we are presented vis- with a flame chart.

  29. 7:03

    Um, and in the flame charts we can see what the CPU is doing while the GPU is not fully utilized. So in, in this case we're, we, we can see that, uh, Python is actually calling, um, eventually the, the CUDA, uh, functions further down.

  30. 7:21

    Um, but oftentimes you, you will see that, like, the CPU is, um, pretty actively, uh, trying to load data and being busy that way and not, not, uh, keeping the GPU busy.

  31. 7:34

    Um, if you are, um, using Python we can see it. If you're using Rust to integrate, uh, with CUDA, for example, um, that also works, but any compiled language is going to show up, uh, in those stack traces, and even some of the interpreted languages are going to show up, uh, like Ruby, Python, um, uh, JVM, et

  32. 7:53

    cetera. So while there's a focus at this conference, um, that we're talking about GPUs here, it really works with, like, any language and, and any application. So web servers, databases, and vector databases for example, um, also are interested in improving their performance obviously.

  33. 8:13

    Uh, something super exciting that we first, uh, uh, introduced this morning, so this is, like, um, super fresh, uh, and hot off the press, is GPU time profiling. So, um, as you heard, like, I, I was talking about, like, these GPU profiles, and we, we look how, how much, uh, time is spent on, uh, individual functions on

  34. 8:35

    the CPU. But we are, like, more in- interested in, uh, GPU time spent, uh, by these functions. So here's, like, a small example of CUDA functions, and basically what we do is we tell the Linux kernel to, um...

  35. 8:52

    Whenever there's a CUDA stack getting put on the, on the CPU, to tell us the start time of that, uh, function, and then eventually tell us the s- uh, time when that, uh, kernel terminates, and then we know the duration of how much time, uh, that particular kernel was spending on, on the GPU.

  36. 9:11

    Um, and, and that's super interesting obviously because now we can actually see how much, uh, GPU time these individual, um, functions are taking on the GPU. And here's a bit m- more of a real [chuckles] world example.

  37. 9:26

    So, um, at the top we can see, uh, uh, on the... Yeah, on the right-hand side we can see, like, the main function in Python, and then calling down into, uh, libcuda down here.

  38. 9:37

    And the width of these, um, stacks that we're seeing is, like, the actual time that we had these functions, uh, take up in, in the GPU. So this is showing CPU and GPU?

  39. 9:50

    Yeah, so this is, like, basically, um, the stack on the CPU- Like- ... down to here, and then the leaf of each stack is the function that was taking time on the GPU.

  40. 10:03

    Okay. W- what do the colors mean? Uh, the colors are different, um,

  41. 10:09

    binaries in this case that are running on your machine. So that's why, like, blue up here for example is, is Python, and then there's, like, some, some I think CUDA, uh, down here.

  42. 10:21

    Yeah, great question. Uh, how do you get started? Because we, um, we run, uh, on Linux using eBPF. You have a binary that you can, that you can run, uh, using systemd or Docker works as well, but we also have a daemon set for Kubernetes.

  43. 10:39

    Um, and you deploy that, you get the m- manifest YAML and give it a token. Um, and then some of our customers are already using it for CPU and memory profiling, and they're starting to also integrate, um, their platforms with our GPU profiling.

  44. 10:55

    Uh, especially like Turbopuffer, um, are, are interested in, in improving their performance of their, uh, vector engine, right?

  45. 11:06

    And that's really it. Um, please visit, visit our booth. Um, um, you get, um... The first 10 people get, like, to sign up for a consultation get two hours for free if you want to, and we can also do discounts for C and CSA startups.

  46. 11:22

    And that's really it. Thank you so much. [outro music]