← All speakers

Bio, Work & Ideas

Nikhil Gupta

Conference affiliation: Crusoe

Nikhil Gupta is a software engineer focused on managed orchestration for GPU infrastructure. In June 2026, he worked as a senior software engineer in Managed Orchestration at Crusoe, bringing practical experience with unified resource pools and cluster recovery to the challenge of keeping demanding AI workloads running.

Gupta advocates Kubernetes for GPU infrastructure. His technical focus connects two needs: researchers need scheduling tools suited to distributed training, while platform engineers need infrastructure that can adapt as resources change and recover when hardware fails. Slurm supplies coordinated scheduling, awareness of how compute resources connect, and familiar job-submission workflows. Kubernetes supplies the operational machinery for managing dynamic resources, monitoring node health, and replacing failed infrastructure. Crusoe’s Managed Slurm combines these capabilities, preserving researchers’ workflows within a platform that operations teams can manage together.

His experience with unified resource pools addresses another practical concern: training and inference demand change over time. Sharing GPU infrastructure allows capacity to move between those workloads rather than remaining divided between separate stacks. His work also encompasses cluster recovery, where replacing a failed node must connect to restoring useful computation. Infrastructure remediation, job scheduling, and restarting training from saved progress are separate tasks that need to cooperate. These concerns place Gupta’s professional focus at the intersection of resource management and workload reliability: making GPU infrastructure useful through changing demand and inevitable hardware failures.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Crusoe combines Slurm’s training scheduler with Kubernetes infrastructure management. The recovery path connects GPU failure detection, node replacement, job requeueing and application checkpoint loading—while researchers keep their familiar commands.

  • Slurm supplies training-aware scheduling and familiar job commands; Kubernetes and AutoClusters connect those jobs to infrastructure telemetry and automatic node replacement.
    1:28 ↗
  • Recovery crosses layers: Slurm cancels and requeues the job, AutoClusters replaces eligible failed nodes, and application code loads the checkpoint.
    10:03 ↗
  • The two-A100-node demonstration reports roughly five minutes for hardware replacement and less than fifteen minutes for end-to-end training recovery.
    11:35 ↗
  • Keeping GPU nodes in Kubernetes allows capacity released by completed Slurm jobs to serve inference workloads.
    15:01 ↗