AI Engineer Code 2025

GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe

Read the talk

GPU Died. Training Didn't: Self-Healing Training at Scale

Crusoe combines Slurm’s training scheduler with Kubernetes infrastructure management. The recovery path connects GPU failure detection, node replacement, job requeueing and application checkpoint loading—while researchers keep their familiar commands.

From a talk by Connor Guerrero, Young Jeong and Nikhil Gupta

At a glance

Ideas worth remembering

  • Slurm supplies training-aware scheduling and familiar job commands; Kubernetes and AutoClusters connect those jobs to infrastructure telemetry and automatic node replacement.

  • Recovery crosses layers: Slurm cancels and requeues the job, AutoClusters replaces eligible failed nodes, and application code loads the checkpoint.

  • The two-A100-node demonstration reports roughly five minutes for hardware replacement and less than fifteen minutes for end-to-end training recovery.

  • Keeping GPU nodes in Kubernetes allows capacity released by completed Slurm jobs to serve inference workloads.

Thousands of GPUs turn hardware failure into routine work

Across thousands of GPUs, hardware failures become an expected part of running training. Fixing each failure by hand makes engineers part of the recovery path, including in the middle of the night. Connor Guerrero, Young Jeong and Nikhil Gupta introduce Crusoe’s response: combine Slurm and Kubernetes so the platform can replace failed hardware and restart training without requiring researchers or platform teams to adopt unfamiliar workflows.

The layers have different jobs. Crusoe Managed Kubernetes, or CMK, manages the underlying cluster. Managed Slurm runs above it, building on Slinky, the Slurm-on-Kubernetes project identified in the talk. Slurm supplies high-performance job scheduling; Kubernetes supplies infrastructure management and observability. AutoClusters adds the hardware-specific response: detect a failed GPU node and replace it.

0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Slurm coordinates training; node recovery adds another job

Slurm’s roots in high-performance computing suit distributed training. A multi-node job needs its ranks—the cooperating processes—to communicate closely across the training network. Scheduling those processes requires attention to both the resources available and their arrangement. Researchers already express these jobs through scripts and familiar sbatch and srun commands.

Several existing Slurm mechanisms serve that work:

  • Gang scheduling: coordinate the resources needed by a job whose ranks must run together.
  • Topology awareness: account for how the resources are connected when placing a training job.
  • Prolog and Epilog: provide hooks for cluster validation around job execution.

The difficulty grows when the same GPU fleet must support training, post-training, evaluation and inference. These workloads change over time, and some may run outside Slurm altogether. Young describes traditional Slurm partitions and resource arrangements as relatively static. Under GPU capacity constraints, that makes shifting hardware toward the workload that currently needs it an operational concern.

Selected presentation frame from GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe at 214 secondsOpen full source frame
A slide contrasts traditional Slurm with autoscaling, health gating, node failures, and observability.

Slurm already has ways to detect bad nodes and requeue jobs. Completing recovery still requires someone or something to connect those capabilities: identify the unhealthy node, drain it, arrange replacement capacity and get the job running again. Prolog checks, open-source tools and custom scripts can do pieces of this work, but maintaining that chain adds burden for the team.

A failed job also needs an explanation. GPU failure is one cause; intermittent network faults or switch problems can degrade performance or stop a job too. Knowing that a job failed does not by itself explain why. The infrastructure needs telemetry that helps distinguish these causes, alongside the scheduler’s view of job state.

2:212:51
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:13 · section reference included

Keep the Slurm interface and share the Kubernetes infrastructure

Kubernetes contributes a mature ecosystem for observability, networking and security, together with service self-healing, load balancing and autoscaling. The organizational problem is that teams often arrive with different tools: inference may already run on Kubernetes, while an incoming training team knows Slurm. Building a separate stack for each preserves familiarity but duplicates infrastructure work and costs.

Crusoe places Slurm inside the existing Kubernetes environment. Its Slurm operator coordinates with CMK and Crusoe’s storage and networking, and manages Slurm users, partitions, configuration and storage. That coordination makes creating a Slurm cluster a single operation rather than a collection of separately configured components.

The same deployment presents two useful views:

  • Researchers: receive an IP address, SSH in with their user and submit jobs to a familiar Slurm cluster. Kubernetes can remain underneath without becoming part of their daily workflow.
  • Platform teams: operate Slurm as another service alongside existing services, using the infrastructure and observability tools they already maintain. Slurm retains its topology-aware scheduling for training jobs.

A shared GPU pool also makes changing demand easier to accommodate. When inference demand rises, hardware can be reallocated from training toward inference. When user load falls, those GPUs can support a larger training run. Nikhil presents this as a capability enabled by the unified stack; the talk does not specify a policy for interrupting active training or deciding which workload wins competing demand.

The shared infrastructure matters for failure recovery too. Managed Slurm can use AutoClusters rather than maintaining a separate hardware-remediation system. The operator connects the scheduler’s job decisions with the platform’s node-replacement decisions.

5:516:21
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:49 · section reference included

From XID 79 to a replacement node and a restarted job

The recovery example starts with an XID 79 error, described here as leaving a GPU completely unusable. The user receives a notification, but no action is required. The Slurm operator marks the affected node down and cancels its running jobs. Before termination, the process receives SIGTERM, with up to two minutes to handle the signal—for example, by saving a checkpoint or flushing logs.

The cancelled job is automatically requeued. AutoClusters then checks whether any pod on the node carries a label that disables automatic replacement. If replacement is allowed, Kubernetes cordons and drains the node: it stops placing new work there and clears the existing workloads. AutoClusters removes the unhealthy node from its pool and replaces it with a healthy node from spare capacity, recording the alert and remediation for the user.

When the healthy node comes online, Slurm starts the job again. Application code loads the model and checkpoint, then resumes training. This division of responsibility is consequential: the platform restores capacity and restarts execution; the application must implement checkpoint saving and loading. A two-minute signal handler provides an opportunity to save state, rather than a guarantee that a failed GPU can still produce a fresh checkpoint.

What connects a hardware alert to resumed training? The flow below separates scheduler actions, node replacement and application recovery. Requeueing preserves the job’s opportunity to run again, but useful training resumes only after both healthy capacity and checkpoint-loading code are ready.

Selected presentation frame from GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe at 614 secondsOpen full source frame
A flow diagram titled “Handling a node failure” shows connected steps across several rows.
How it fits togetherHardware recovery and training recovery meet at job restart

Notify the user that a GPU is unusable.

Automatic replacement depends on the pod-label check and spare capacity. The application restores training state after Slurm restarts the job.

10:0310:33
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:03 · section reference included

The demo’s utilization drop becomes a checkpoint restart

The demonstration follows a PyTorch training script launched with sbatch against a pool of two A100 nodes. Connor then runs a script that tells Crusoe’s internal monitoring system an XID 79 error has occurred on one GPU. This injects the failure signal into the remediation path; it does not establish that the demo physically damaged a GPU.

GPU utilization drops, and node replacement begins immediately. Connor reports roughly five minutes from detecting the critical hardware error to returning a healthy node to the pool. The remaining time is mainly application-level work. That distinction explains why replacing hardware and recovering a training run have different completion times.

The job restarts, loads its checkpoint and continues training. Utilization returns to normal, with reported end-to-end downtime of less than fifteen minutes in this example. The alerts inform the user throughout, without asking them to perform recovery steps. The timing belongs to this demonstration, rather than a recovery-time guarantee for every model or cluster.

11:3512:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:30 · section reference included

Provision the recovery path, then keep both systems in sync

One-Click Slurm packages the environment into a provisioning operation. Despite the name, the interface described is a command: one command provisions the Kubernetes cluster, Slurm controller, login nodes and storage. A second command adds the GPU node pool. Researchers can then SSH in and submit jobs, with AutoClusters enabled by default.

Selected presentation frame from GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe at 829 secondsOpen full source frame
A slide titled “1-click Slurm” shows two command examples for provisioning.

The closing architectural point is synchronization. Traditionally, Slurm and the infrastructure beneath it can be operated as separate systems. Here, a failed Kubernetes node can be cordoned automatically and that signal propagated to the Slurm operator. The scheduler and infrastructure therefore respond to the same failure, rather than leaving an engineer to reconcile their states.

That connection preserves the two teams’ working interfaces. Researchers keep their Slurm commands without having to SSH into pods or manage containers. Kubernetes teams keep the telemetry and observability they expect. The integration absorbs the coordination work beneath those interfaces.

The GPU nodes are “Kubernetes nodes first.” When a Slurm job finishes, Kubernetes can schedule inference pods on those nodes instead of leaving them idle while they wait for another training job. This is the concrete resource-sharing case at the ending: completed training releases useful capacity for a different kind of workload.

Designing for inevitable failure means arranging the whole response in advance: detect the error, stop scheduling onto broken hardware, replace the node, restart the job and restore application state. Automating that sequence lets engineers spend their attention on the application instead of becoming the mechanism that gets training running again.

13:1813:48
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:18 · section reference included

Read the complete timestamped transcript
  1. 0:12

    All right. Hello, everyone. Uh, my name's Connor. I'm a developer advocate at Crusoe, and I'm joined by my colleagues Nikhil and Young. And we're gonna be talking today about the infrastructure we've built at Crusoe, and how it handles critical hardware errors. And so our customers, when they're running extremely large scale training workloads across thousands of GPUs, GPU failures are inevitable. And so at scale, manual remediation is completely unsustainable. And so what we'll cover today is how

  2. 0:42

    by using both Slurm and Kubernetes, we can provide a robust platform with automated remediation that maintains ease of use for both machine learning engineers running jobs and the platform teams managing the underlying clusters.

  3. 0:59

    So to provide a little bit more context on what Crusoe Cloud is, for those of you who haven't heard of us, um, we provide infrastructure as a service consisting of compute, storage, networking, and the latest gen GPUs from NVIDIA and AMD. But the focus of this talk will be on the tooling that we've built on top of that to orchestrate these large fleets of GPUs, including our managed Kubernetes service, called CMK, and our managed Slurm service, which is built

  4. 1:28

    on top of CMK and built off of Slinky, which is SchedMD's official open source project for running Slurm on top of Kubernetes. And this allows us to get the high performance job scheduling of Slurm with the infrastructure resilience features and observability of Kubernetes, including one of the systems we've developed called AutoClusters, which is our system for automatically detecting and replacing failed GPU nodes when a GPU failure has been detected. And so

  5. 1:58

    to explain how and why we landed on this architecture, it's helpful to understand where traditional Slurm both excels and falls short, which Young will get into now.

  6. 2:13

    Thanks, Connor. Um, traditionally, I mean, looking at it from the trad- traditional perspective...

  7. 2:21

    Sorry, can, can you hear me now? Okay, cool. I'll, I'll just do it like this. But basically, traditionally, Slurm was built, Slurm was built by the researchers for the researchers, uh, universities, labs, over twenty years ago, uh, built for a high performance computing, uh, which as we've observed with our customers and others that are using Slurm today, translates very well to modern AI, specifically around training workloads. This idea that when you're running a multi-node training

  8. 2:51

    job, you need this idea of tight collective communication across the ranks that are-- that you're using, uh, including, you know, specific network built for training workloads today. Um, concepts like gang scheduling, topology awareness, uh, Prolog and Epilog for cluster validations. These are all the features that are readily available with Slurm today that helps researchers run their jobs, uh, essentially running

  9. 3:21

    their jobs with tools that are already familiar to them, scripts, sbatch, srun commands, the things that are pretty standard in your Slurm cluster today. Now, where we feel like this is a bit of a shortcoming or a short-- falls a little bit short around the modern training is the things that you see on the screen right now. Um, AI training workloads today are very dynamic, and it doesn't just include training. Uh, there's a lot of other aspects that your AI

  10. 3:51

    labs and oth- other companies are running as part of your, your AI workloads, including training, post-training, eval, inferences, some of which may or may not even land on your Slurm cluster today. But because of, of what we've observed as, as a GPU capacity constraints and other things, kinda limits you to how you need to be able to be more dynamic around the resources that you have. Uh, Slurm traditionally is pretty static, and it's partitioning in other, other, uh, functionalities. The

  11. 4:21

    other thing I wanna mention is this idea of health checks and maintaining your nodes. Uh, you know, you can make sure your nodes are healthy or not healthy. GPUs do fail. How do you manage that? Uh, Slurm obviously has features around it to be able to, you know, re-queue the jobs, uh, detect bad nodes, uh, using things like Prolog and other open source tools that are now available to be able to do this. It just adds an operational burden, uh, that we've seen from customers

  12. 4:51

    today. You know, this idea that you have to detect if nodes are bad, you have to drain the nodes when they're bad, you have to kind of work around that. That could either be manual or you need some sort of an automation or scripting to be able to do all of this. And the last thing is observability. Um, you know, jobs can fail. Jobs can also come with degraded performances for any number of reasons. I mean, GPU failure is just one of them. There's other things like network flaps and other things within your, your switches

  13. 5:21

    that may be causing degraded or failed performance. Uh, Slurm knows when a job fails. It doesn't necessarily be able to tell you everything about why it failed, and that's where the observability comes into key, uh, key play. So all of this kind of encompasses the reason why we, we built this product that Nikhil's gonna go over and sort of describe some of the specifics around Managed Slurm.

  14. 5:49

    Um,

  15. 5:51

    all right. Uh, thanks, Young. So yeah, I'm gonna be talking about our, uh, Crusoe Managed Slurm product. And, uh, like, like we mentioned before, um, there's a lot of benefits, uh, that Slurm can bring to the, uh, training teams and, uh, researchers that are building the next bes- best models. But there are some issues with, you know, setting up Slurm, managing it, keeping it reliable. Uh, and on the other hand, you have Kubernetes, which

  16. 6:21

    is the OS of the cloud. It's wide-widespread use and, uh, familiarity. It's a very mature ecosystem with lots of tools, extensions, and plugins for whatever you need to add, whether it be observability, networking, security. It helps, uh, maintain services with their, you know, built-in self-healing, uh, load balancing and auto-scaling that keeps everything running smoothly. And so we often see that, uh, customers will, uh, come with us starting with, say,

  17. 6:51

    inference running on Kubernetes, and then they wanna expand to training, uh, or vice versa. And so maybe a Kubernetes-familiar team struggles with deploying and managing a Slurm cluster, um, or, uh, the incoming Slurm native training team is unfamiliar with the existing Kubernetes frameworks, uh, and adds, you know, additional friction that just slows down your team. And with, you know, how fast everyone's trying to get to, uh, the next best model and

  18. 7:21

    next best product, um, teams get set up with the tools they're familiar with, and you end up with two separate infrastructure stacks, which has its own host of extra operational burdens, uh, and extra costs. So we decided to build out our Crusoe-managed Slurm, uh, on top of Kubernetes. So we have our Crusoe Slurm operator, uh, which coordinates with our Crusoe-managed Kubernetes platform and the broader Crusoe cloud, so including storage, networking, et cetera.

  19. 7:51

    And, uh, CSO manages, uh, Slurm users, Slurm partitions, configuration, and storage, so you don't have to do it piecemeal. You just, you know, one, one command, you create your Slurm cluster, and there it is existing in your Kubernetes cluster. So the benefit is for the Slurm users, um, it's just another Slurm cluster. They just need, uh, an IP address. They can SSH with their, uh, user, and now to them, it's a Slurm cluster with all the bells and whistles that they need. They don't even need to know that Kubernetes is running

  20. 8:21

    underneath. For your platform team and your, uh, infrastructure team, uh, the benefit is Slurm is now just another service that plugs into your existing infrastructure. You don't need to set up an entirely new infrastructure stack, new observability tools, new, you know, on-call teams, new runbooks, nothing like that. It's all built into the same thing. It runs alongside all your other services. And, um, resources in Slurm, so any jobs, make use of the, uh, built-in topology-aware scheduling,

  21. 8:52

    um, and the platform team just sees it as another service. Um, and now since all your GPUs and hardware are in one cluster, they can move around seamlessly. Uh, it enables, uh, more interesting, uh, tooling and dynamic a-availability, like when you need to burst your, uh, inference service when you're reaching, uh, high, uh, users, high number of users. You can take away some of your-- or even reallocate some of your hardware from your training cluster to your inference service, or vice versa.

  22. 9:22

    You know, when your user, uh, load goes down, you have all these GPUs sitting idle. That's the perfect time to up your training job or do a, a larger hero run across more GPUs. Um, this is something that's only possible when you have everything in a single, uh, infrastructure stack like, uh, Crusoe-managed Slurm. Um, an additional benefit is, uh, now Crusoe-managed Slurm can make use of our AutoClusters product, which is our auto-remediation for additional hardware errors.

  23. 9:52

    So Connor is, uh, gonna walk through that and show a demo of, uh, of how that works.

  24. 10:03

    So this is an overview of what exactly happens when AutoClusters detects a critical hardware error, in this case, an XID seventy-nine error, where one of the GPUs is completely unusable. And so the first thing that happens is that the user is notified, although there's no action required from them. Then the Slurm operator sets that, that specific node to down and cancels any running jobs. There's then a, um... The process rec-receives a sigterm signal before

  25. 10:33

    being terminated, and you have up to two minutes to handle that sigterm to either sa-save your checkpoints or flush your logs before that node is eventually replaced. The cancel job is automatically re-queued, and then AutoClusters starts. So first, it verifies that no pods on the node have a label that disables automatic replacement. Assuming that node can be replaced, the Kubernetes node is then cordoned and drained, and then the

  26. 11:03

    unhealthy node is removed from the node pool and replaced with a new healthy node from spare capacity. And then a record of the alert and the remediation is logged for the user's reference, uh, but there's still no action required from them. And then once that healthy node comes back online, then the Slurm job starts, your application code runs to load your model and lo-load your checkpoint, and resume training.

  27. 11:30

    And so here's what that looks like, um,

  28. 11:35

    in a demo. And so here on the left, I'm just using sbatch to launch a PyTorch training script. On the right is the Crusoe cloud console, where I have a GPU node pool of two A100 nodes. And I'm gonna launch this as normal, um, and then what I'm gonna do is run a script in-- to tell our back-end internal monitoring system that there's been an XID seventy-nine error on one of the GPUs, and this will trigger

  29. 12:05

    AutoClusters. And so we can see immediately the GPU utilization drops, and so we know that that n-- that GPU has gone offline, and AutoClusters node replacement starts immediately. And the full process from detecting a critical hardware error to getting a healthy node back into the node pool takes roughly five minutes with AutoClusters. The rest of the time is mainly application-level code. And so in this example, it's a

  30. 12:35

    little under fifteen minutes, um, from end to end. And here I'm just showing some of the alerts that, uh, the user receives when there's a critical hardware error like this. And so a-as you'll see, um, there's no action required from the user. The Slurm job will restart. It'll load the checkpoint and continue training where it left off. And we can see the GPU utilization returns back to normal, and the total downtime is less than fifteen minutes. And this is, um, much better

  31. 13:05

    than having an engineer have to log on in the, in the middle of the night and spend hours debugging something. All of this is handled automatically on our platform.

  32. 13:18

    Now, historically, creating this type of environment has been challenging to merge both Slurm and Kubernetes. And so we're really excited to announce what we call One-Click Slurm, where everything that we've talked about in this presentation, this entire environment can be provisioned with, with just a single command. And so this one command includes provisioning the underlying Kubernetes cluster, the Slurm controller, login nodes, and storage. And then adding a GPU node

  33. 13:48

    pool is only one command beyond that. And so this means that engineers can simply SSH in and start running jobs immediately, and AutoClusters is, is enabled by default.

  34. 14:02

    So wrapping up what we covered, um, starting with the limitations of traditional Slurm, it gives you the high-performance job scheduling, but not the automated, the automated healing that modern AI infrastructure requires. And so by building Slurm on top of Kubernetes, we can enable these advance, advanced features without having either team change, uh, their workflows. Then we covered the full node remediation flow from an error being detected to the Slurm job resuming and the

  35. 14:31

    checkpoint-- and the training resuming. And so the key takeaways here, first, is that traditionally, Slurm and the infrastructure that it runs on are two separate systems. But together with Kubernetes, both systems stay in sync. And so a failed GPU node can get cordoned automatically, and that signal will propagate str- directly to the Slurm operator. Also, like I mentioned earlier, neither

  36. 15:01

    the platform teams nor the machine learning teams have to change their workloads. Researchers can still use the same commands to launch Slurm, and although Kubernetes is running underneath, researchers never have to SSH into a pod or worry about containers. But the Kubernetes teams still get access to the same level of telemetry and observability that they're used to. And then lastly, like Nikhil mentioned, the GPU nodes are Kubernetes

  37. 15:31

    nodes first, and that means that Slurm is just a workload running on top. And when a Slurm job finishes, those GPU nodes don't have to sit idle waiting for another job. You can schedule inference pods on those immediately through Kubernetes because from the platform's perspective, they're still just nodes in a pool.

  38. 15:55

    And so the, the core idea here is that failures are inevitable, um, and so in today's AI landscape, the architecture of infrastructure should be designed in a way so that when there is a critical error, all the right actions are handled autonomously, and engineers can focus on building their application and not worrying about the infrastructure.

  39. 16:20

    So hopefully, this was an informative talk. Um, if you have any questions, feel free to, um, connect with us at Booth UG-12 or on social at crusodev. And, uh, thank you for listening, and feel free to ask any questions. Thanks.