AI Engineer World's Fair 2026
Infra behind Krea 2 - How to train and serve at scale
About this talk
Gabriel Jorge Menezes explains how Krea trained its Krea 2 image model from scratch across large InfiniBand-connected GPU clusters, adapting ideas from language-model research to diffusion transformers. He details operational observability, GPU thermal throttling, misleading utilization figures, tensor-core and networking metrics, then describes Kueue gang scheduling, Kubernetes workload priorities, and practical infrastructure for serving accelerated image-generation models.
Chapters
- 0:00Introducing Krea 2 and its open-source Turbo model
- 1:49Scaling diffusion-transformer training across GPU clusters
- 4:12GPU, tensor-core, InfiniBand, and NVLink observability
- 8:43Kueue gang scheduling and Kubernetes workload priorities
- 10:27Inference operations and closing remarks
Talk transcript
- 0:00
[upbeat music] Hello, everyone. Uh, my name is Gabriel.
- 0:15
Uh, I work at Krea, and I'll be talking about the infrastructure that allowed us to train K2, and also how we serve it. Um, so what is K2? K2 is our pre-trained from scratch model, uh, we just released like less than a month ago.
- 0:29
And the whole idea about training this model was because we are kind of bored of AI images. They are quite, you know, saltless. They have no spice. And the whole idea was we want to give creatives tools to, to explore out of distribution extremely interesting images, do composition, and like actually give tools to creatives, and that was
- 0:50
the whole idea of the model. Uh, the model was trained from scratch, uh, no base checkpoint, not anything. Everything done in-house, and also, as said, built for exploration. And right now, uh, this is what you can get out of Krea 2, very different styles and, you know, like pixel art and, and like, like photoreal and like some
- 1:12
silly stuff. Uh, and whole idea of the model, as said, let's explore this medium. Uh, Krea 2 is open source. Uh, right now, uh, you can go play with it.
- 1:23
There is two checkpoints we also serve in production. Uh, there is a raw checkpoint, which is pre-trained, so people can post-train and do whatever they wish to do with it.
- 1:32
Uh, and there is also the post-trained version, which is the Turbo one, which is very, very fast. You can get like an image in like, I don't know, less than a second.
- 1:38
And this is-- Yeah, so like this type of images you can get in less than a second. Uh, on Hugging Face GitHub, go, go play with it. Also, you can go in production on Krea.ai and go play with it.
- 1:49
Uh, so let's talk about how we train this model. Uh, I said it's gonna be how we train and how we serve. Uh, first, the model was trained from scratch on thousands of GPUs.
- 1:58
Uh, we have a big cluster, one main cluster with a lot of GPUs, all InfiniBand connected, and you put those GPUs work, and it train. But like that, that's-- I wish it was that simple, but it's not.
- 2:10
Uh, so at the beginning, we did a bunch of like small ablations on like,
- 2:14
like a small number of GPUs to see like how things go, right? So you wanna test some hypothesis, and you do a small number of GPUs and let it train for a little bit.
- 2:22
Oh, this works. It doesn't work. Let's scale. Uh, and as we are like training this model, the whole idea was to like kind of bridge the gap between like LLM research and diffusion transformers.
- 2:34
Uh, so my AI researchers, they ported a lot of research from LLMs into, into DiTs, and the whole like in-- like the whole like architecture of the model was meant to be extremely, extremely simple.
- 2:46
And so like it is very, very dumb, but like very effective. Uh, and so let's start talking about numbers. Uh, incredibly, our maybe skill issue on our part, maybe our cluster, uh, was very interesting as we were like scaling.
- 3:02
Uh, when you do like small experiments, experiments would like run for days and like even less than we would like to, but like they would still run fine. Uh, and as we like start scaling, like getting like more and more and more GPUs, like one twenty-eight, two fifty-six, five twelve, whatever number, and like you scale and scale,
- 3:18
like things start crashing more. That's expected, right? Like there's more surface area for things to break, and things gonna go wrong. And a lot of the times things would go wrong in silent ways.
- 3:28
I know, Nicotimeout, like it just crashes, and like the metrics are all good. Uh, and it is extremely annoying stuff. At the beginning, we were like paranoid, swap node, change nodes, whatever, whatever, whatever.
- 3:40
And we learned that like sometimes you just let it crash. It crashes, like it runs for like an hour, crash, runs for an hour, crash, and then it runs again on the same set of machines, same code, same data for like twelve hours, sixteen hours, twenty-four hours.
- 3:52
Uh, there is this paper from Meta that kinda gives you like a rough estimate of like how many failures you should expect. Uh, we, we kind of saw the same pattern, but like not the s-the same numbers.
- 4:02
Uh, our runs would last way, way less than this. Uh, so it was extremely annoying, and you can imagine doing large-scale pre-training on runs that last less than eight hours, it is a problem, right?
- 4:12
You want-- kept those GPUs fed, and if things are crashing, you're not doing progress and losing time, and model's gonna be late. Uh, so for us, extremely important, and like for-- at least from the infra side, was to get metrics.
- 4:26
Metrics are everything. That's how I can support my researchers. That's how I have visibility in the system. And like if you're doing large-scale pre-training, I highly, highly recommend for you to invest heavily on metrics.
- 4:37
Uh, don't, don't go blind, because you're gonna go, you're gonna go crazy. So I'm gonna share some of the metrics that like were important for us, and like they're quite silly, but like extremely ef-effective.
- 4:46
Uh, first one, like GPU temperature. Uh, GPUs are very, very annoying. They-- If you're-- If you have a single GPU that is like a bit warmer than the others, it's gonna start like throttling and slow down when-- and the training is gonna be unstable and have weird problems.
- 5:00
So for us, it was like if there is any GPUs above like seventy-eight degrees, you, you remove them. Don't, don't think about it. Don't try to fix. Don't, don't try to be smart.
- 5:08
Ju-just remove the GPU. Uh, you're gonna save time and just ask your provider like, "This GPU is hot. Uh, please replace it." Uh, and this-- There, there is two metrics that at beginning we did not fully understand, and as we are like getting more and more used to GPUs and like how they work, uh, there is GPU
- 5:27
utilization, which is a lie. Uh, don't trust this. Uh, this is dumb. Uh, this tells you, "Oh, the GPU is doing work," and this is amount of time GPU doing work, but like not good work, not how efficient the GPU is working.
- 5:38
So like as you can see during our pre-training, the GPU is at a hundred percent. This is not true. We are not fully utilizing the GPU. Uh, this is a hundred percent a lie.
- 5:46
What we would use a-as a proxy was tensor core utilization. This is actually how, how much of a tensor cores you're using and like how effective they're being. Uh, and it was also very interesting as we're doing pre-training and then you go from pre-training to mid-training to post-training, you start like scaling the-- like the resolution of the
- 6:02
images. So like for example, pre-training, we do like, I think one twenty-eight, two fifty-six, five twelve
- 6:07
1024 pixels as they scale. And you could see the tensor core utilization go up as we, like, would scale on these, these resolutions, because now we're doing more work on, on, on images.
- 6:17
And like, also very, very interesting was InfiniBand and NVLink metrics. These by default not exported by, uh, the NVIDIA metrics, the DCGM stuff. Some NVLink stuff, yes, but no InfiniBand.
- 6:29
So if you don't have InfiniBand metrics, uh, go, go, go get it. Uh, I'm telling you right now, if you're doing large-scale pre-training with a bunch of GPUs talking to each other between machines, and you have no InfiniBand metrics, you are doing something wrong.
- 6:40
Uh, this was probably, like, the most important stuff for us, because most of our failures were, like, related to, like, cross-node communication. So, like, for example, here you just have throughput, but, like, on our, like, refined dashboard, we have a bunch of stuff.
- 6:52
Like, from, like, wait, like, when y- when a message is sent on a, on a, on a fabric, like, how much time the, the, the message is waiting, or, like, the, like, number of errors and different types of errors and, like, number of packets, and all of the things that it, it...
- 7:06
like InfiniBand is exported, uh, we collect. We had to build custom stuff to get this. It was not hard. Uh, you can figure it out. It's very, very easy.
- 7:13
Same thing for NVLink. Uh, NVIDIA exports some stuff about NVLink, but, like, for example, NVLink, NVLink errors, NVIDIA doesn't export this. Uh, so you can collect this, and like, as I said, InfiniBand was extremely important.
- 7:25
NVLink was a little bit less. Uh, in some cases, this helped us catch some problems in, like, especially 'cause NVLink is single node, right? Like, the communication inside a node.
- 7:33
So sometimes a single node would have a weird failure where the GPU seemed to be fine, but, like, some weird error happens, and then you can look at NVLink, NVLink errors, and, like, you see all errors happening.
- 7:43
Uh, and then replace that machine. So go get these metrics. They, they're extremely important, and without this, we would not be able to, to train, uh, at all. Uh, also, as I said, our, our trainings would crash constantly.
- 7:57
Uh, and a hacky way to do it, to fix the problem, just checkpoint. Uh, use and abuse the file system that you have. At the beginning, we used Ceph.
- 8:05
Ceph did not work well. Uh, was very annoying. It broke. We would not trust the data. So I recommend if, if you have the money, go, go with something paid, uh, th- because you can trust your data.
- 8:17
You can see numbers. This is our, our Weka cluster. We can do, like, 1.2 terabytes of se- uh, second of reads, almost a terabyte of re- uh, of writes.
- 8:25
Uh, and the file system would not choke on the training. So we could checkpoint every, like, 30 minutes, 20 minutes, produce, like, a terabyte of data in less, less than 30 seconds.
- 8:33
Uh, so this would not delay trainings. That was, like, probably one of the, like, most important things that we did to, to, like, recoup the loss, like, just checkpoint.
- 8:43
Uh, don't, don't, don't think about it. [laughs] And, like, how we serve. Uh, this goes in connection on, on how the trainings are launched. Uh, because at the beginning, I don't want my researchers to think about GPUs.
- 8:55
I just want them to launch stuff, and this goes into a queue, and if we have GPUs, we have GPUs. If we don't have GPUs, we don't have GPUs.
- 9:03
Uh, so Kueue, uh, this is a op- open source project. You can look it up. It does Gang scheduling. Uh, Gang scheduling extremely important for, for, for trainings in general.
- 9:13
Uh, and this gives us a semantic of, like, two tiers of priority, uh, where you have a workload priority, and this, you can say, like, "Oh, this training is more important than this one," so it keeps on the queue, in front of the queue.
- 9:25
Uh, and then a- after this, we have the normal Kubernetes priority if we're used to Kubernetes. And the way the system works is, like, the training pods, they always have, like, high priority for everything.
- 9:36
So, like, once they are admitted, they're gonna schedule. If there is inference running on those machines, the inference gets kicked out, and you would say, "Oh, this is bad.
- 9:42
Production is gonna go down." No, you can build on top of that to, to make production not go down, which is very, very cool. Uh, the on- one of the problems with Kueue, which is annoying, you can automate that.
- 9:51
We have not. It's just that you specify the queues. The queues have, like, amount of resources, CPU, and, like, NVIDIA GPUs, memory, whatever. Uh, but this is manually, like, manually, like, specified, and at least our cluster is quite fluid.
- 10:07
Nodes phasing in and out of existence. They go into, to maintenance, whatever. You lose a few nodes here and there. Uh, this number gets out of sync, and sometimes these, these would break Gang scheduling.
- 10:17
Uh, so FYI, th- this is a bit annoying. You're gonna face this if you use Kueue. Uh, but yeah, very good project. Kubernetes 1.35, uh, we have not had the chance to play with it.
- 10:27
Has Gang scheduling, uh, out of the box. Something very similar to Kueue. Uh, so maybe you can use Kubernetes 1.35. Uh, and as I said, this, this is the system that we built that allowed us to train using the whole cluster.
- 10:39
As I said, we have one big cluster that runs production and, and trainings. Uh, so I don't want my researchers to think about GPUs, and I don't want make the trainings and research be delayed because production is running, right?
- 10:52
Uh, production is l- lower priority. The site still needs to work. People still need to be use the website. But, like, the GPUs, they're, like, the value that we get out of the GPUs during training is more, like, higher than we get out of production.
- 11:04
So the whole system works by default, where there is this magical system that I'm gonna explain a l- in a bit, that allows us to flip traffic between clusters magically, and not just clusters, uh, like, external providers, GPU rentals, whatever.
- 11:16
And you can see, like, the green, like, the, the dark green is, like, inference running in cluster. Uh, and then someone launches the train, and then suddenly start flipping to the other cluster, and then training is done or whatever happens, it flips back, so we stop wasting money, and this is seamless.
- 11:33
No one needs to think about it. Uh, the whole system, like, like, handles itself, and you get this, this very nice pattern of, like, I'm gonna use all the GPUs in my cluster, uh, for trainings.
- 11:45
Production is gonna run somewhere else. I don't need to think about it. My users on production, they're not gonna feel anything. Uh, research is gonna be happy, and we can get values out of the GPUs.
- 11:54
Uh, so how does this work? There is this very nice project called Virtual Kubelet, also open source. Uh, you can build on top of it. It is a very nice code base.
- 12:02
Uh, and this works by creating a fake machine in Kubernetes. Uh, Kubernetes has nodes. This creates a fake machine that is, like, up to you to control how it works.
- 12:10
So Kubernetes does normal scheduling, as you would expect. Things would go into, into these nodes. For example, here, all the GPUs in the cluster are in use, right? So this pod goes into the Virtual Kubelet node, and in there you can do whatever.
- 12:23
This is the system that we built. There is, like, you receive the pod spec, and then you find a provider. You can, like, this is up to you. Let's say you have a deal with some provider that gives you nice prices, you integrate, integrate into here.
- 12:34
We built, like, some nice interfaces to b- to, to, like, not leak things. So, like, we just implement a provider, and there is a, a, a algo that decides which one it goes.
- 12:44
Uh, you translate the spec of the pod into, into the provider stuff, and you deploy, and then you have something that reconciles between both sides. And what's extremely, extremely nice, if you guys know about Kubernetes, Kubernetes has, like, the horizontal pod autoscaler, which, like, scales the number of, of replicas.
- 12:59
Uh, number of replicas, uh... Could you stop being annoying? Thank you, sir. I appreciate it. There you go. Let's go back. Uh, now we, uh, go back. Ah, back, back.
- 13:12
There we go. Uh, like, Kubernetes has the HPA, and the HPA scales the pods. And so if something fails, it is very interesting, you don't need to handle the fail.
- 13:21
Uh, the only thing you need to handle is like, oh, something failed. You mark the pod as failed. Kubernetes is gonna detect that something has failed and creates a new one.
- 13:27
Uh, you don't need to try to save the world. Let Kubernetes handle it for you, which is extremely nice way of handling stuff. If something breaks on your side, something breaks on the, the other side, just mark as failed.
- 13:37
Let Kubernetes handle, create a new one, and things keep working. Very, very nice way to handle stuff. Uh, and also very interesting, let's say you have GPUs on your cluster available, right?
- 13:45
You don't wanna waste money. This would be very, very bad. So the system works about, like, works with, like, using Taints, Kubernetes Taints. They allow and disallow things to run.
- 13:55
Pods have tolerations for the Taints. And when we have GPUs in the cluster, uh, if you look back, there is the Taint system in the bottom. The Taint system, it is what would, like, by itself decide what y- like, if we have GPUs or not GPUs in the cluster.
- 14:09
Uh, and this adds a Taint into the, into the node when we have a lot of GPUs. So, like, a lot of GPUs in the cluster, we Taint the node.
- 14:16
Nothing can schedule on it. Uh, so we stop wasting GPUs. G- like, the pods, they go into, into the GPUs in the cluster. We don't waste money. Uh, and then imagine someone launches a training, right?
- 14:27
This training is gonna take all of the GPUs in the cluster. It's gonna hog all of the GPUs. No GPUs in the cluster. The system detects this, removes the Taint, new pod scheduled there.
- 14:34
Very, very nice. Uh, you also don't think about it. Uh, and for us, for example, we use just some Prometheus metrics. That's how we do it. Very simple, uh, but it works very, very, very well.
- 14:44
Uh, you let the system run by itself. Someone's gonna launch stuff. You're gonna, like... The training itself is gonna kick off the pods. It's gonna schedule. It's gonna take the GPUs.
- 14:53
The system's gonna detect that, remove the Taint, pod scheduled there. Very nice. Someone finished the training. Now we have pods running on the other side. You're wasting money. Uh, how do we fix this?
- 15:02
Uh, same thing. You run something else that detects the system and removes things back. Uh, in this case, the deschedule. Uh, once the Taint's added back, so, like, GPU's available, we add the Taint.
- 15:12
The deschedule sees, oh, these pods, they don't tolerate the Taint. I'm gonna migrate them back. And you can ask, "Oh, why you don't use a no, no, no execute Taint?"
- 15:20
No execute in Kubernetes would kick everything out at the same time. So as, as, like, the moment you put the Taint, everything would be kicked out, and that's bad.
- 15:27
Production would go down. So, like, this system mi- like, slowly migrates the pods back so production does, doesn't go down, and we don't waste money. Um, it is, like, a very, like, self-healing system.
- 15:38
You don't need to interfere with it. Uh, it just runs. Uh, yes, of course there was bugs in the beginning, and nothing's perfect, but, like, once you calibrate it, was, was, like, very, very well and, like, changed the way we do research.
- 15:49
'Cause no one else needs to care about GPUs. They just launch stuff. If we have GPUs, we have GPUs. If not, we don't have GPUs, go into the queue, and we fully utilize the cluster for trainings.
- 15:58
Production run somewhere else, and the GPUs are doing, like, useful work. And also, like, if you're doing, like, diffusion transformers, they're not huge like LLMs. They need, like, multi-node, like, inference.
- 16:08
Uh, something that we learn, like, whatever GPU works. Uh, the GPU can be hot, falling out of the bus. It can be exploding. Uh, inference still gonna run. It is very interesting.
- 16:17
So, like, you can have very, very bad GPUs for inference, uh, and everyone's gonna be happy. Uh, we are hiring. Uh, if you're interested in building this sort of stuff, uh, doing large-scale pre-training, uh, build this sort of system for researchers, shoot me a message at [REDACTED:email_address].
- 16:33
There is also jobs listing, and that's it. Thank you. [audience applauding] [upbeat music]