← All speakers

Bio, Work & Ideas

Young Jeong

Conference affiliation: Crusoe

On this page

Young Jeong’s work connects Kubernetes operations with the performance and reliability of large-scale AI workloads. As a Staff Solutions Engineer at Crusoe in his collaborative work on GB200 fine-tuning, he helped explain how GPU placement, memory constraints, and communication between accelerators shape a training system’s behavior.

From application modernization to AI infrastructure

At AWS, Jeong worked as a partner solutions architect focused on application modernization. His collaborative writing addressed the transition from working application code to systems that teams could deploy, debug, and operate reliably.

In 2021, he co-authored a guide to machine-learning workflows with Kubeflow and D2iQ Kaptain. Deploying model code was only part of the problem: data scientists also needed reproducible environments, distributed training, experiment management, security, and a route into production. The guide showed how a Python interface could run parallel tuning experiments, reducing the need to switch between notebooks, container builds, Kubernetes manifests, and command-line tools.

His subsequent writing extended that operational focus to migration and debugging. With Rosemary Wang and Welly Siauw, he co-authored a 2023 Consul service-mesh implementation that let services move gradually from Amazon ECS to Lambda while maintaining controlled communication across both environments. With Eran Kinsbruner and Shakthi Dakuri, he developed a guide to live troubleshooting on Amazon EKS, addressing developers’ difficulty understanding application behavior after deployment to remote clusters.

His work at Crusoe applies that operational perspective to distributed AI, where GPU memory, high-speed interconnects, scheduler placement, and communication libraries must work together. In the Llama 3.1 fine-tuning collaboration with Zoom engineers Jack Jin and Han Jin, executable Kubernetes configuration accompanies workload comparisons on GB200 NVL72. The work explains why adopting new GPUs requires attention to the whole training system, including job placement and how model state fits into memory.

Making GPU capacity useful

  • Topology-aware scheduling: The GB200 collaboration connects Kubernetes Dynamic Resource Allocation with rack-scale NVLink domains. Training pods need placement that respects the fast connections among GPUs; communication paths and parallelization strategy help determine how much performance a workload gets from the hardware.
  • Debugging distributed training: Jeong’s NCCL troubleshooting guidance explains how a stack-limit setting can prevent a large GPU job from starting. NCCL, the communication library used between GPUs, performs a recursive hardware-topology search that can require more stack space than the process receives. An unlimited setting can paradoxically leave too little. The guidance specifies an eight-megabyte stack limit inside the Slurm job, applying the fix where training actually runs.
  • Inference topology tuning: Jeong and Alex Akesson co-authored Crusoe’s MLPerf Inference v6.1 results in September 2026. For the tested Qwen3-VL workload on 16 GPUs, changing the allocation between prefill—the processing of an input prompt—and decode—the generation of subsequent tokens—from 75:25 to 50:50 increased Interactive throughput from 15.85 to 39.85 queries per second. The result illustrates how the serving arrangement can change performance without changing the hardware.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Crusoe combines Slurm’s training scheduler with Kubernetes infrastructure management. The recovery path connects GPU failure detection, node replacement, job requeueing and application checkpoint loading—while researchers keep their familiar commands.

  • Slurm supplies training-aware scheduling and familiar job commands; Kubernetes and AutoClusters connect those jobs to infrastructure telemetry and automatic node replacement.
    1:28 ↗
  • Recovery crosses layers: Slurm cancels and requeues the job, AutoClusters replaces eligible failed nodes, and application code loads the checkpoint.
    10:03 ↗
  • The two-A100-node demonstration reports roughly five minutes for hardware replacement and less than fifteen minutes for end-to-end training recovery.
    11:35 ↗
  • Keeping GPU nodes in Kubernetes allows capacity released by completed Slurm jobs to serve inference workloads.
    15:01 ↗