▶ Watch ↗AI Engineer Code 202516:54
GPU Died. Training Didn't: Self-Healing Training at Scale — Crusoe
Read the full talk →Key ideas
Scroll to read ↓Crusoe combines Slurm’s training scheduler with Kubernetes infrastructure management. The recovery path connects GPU failure detection, node replacement, job requeueing and application checkpoint loading—while researchers keep their familiar commands.
- Slurm supplies training-aware scheduling and familiar job commands; Kubernetes and AutoClusters connect those jobs to infrastructure telemetry and automatic node replacement.1:28 ↗
- Recovery crosses layers: Slurm cancels and requeues the job, AutoClusters replaces eligible failed nodes, and application code loads the checkpoint.10:03 ↗
- The two-A100-node demonstration reports roughly five minutes for hardware replacement and less than fifteen minutes for end-to-end training recovery.11:35 ↗
- Keeping GPU nodes in Kubernetes allows capacity released by completed Slurm jobs to serve inference workloads.15:01 ↗