← All speakers

Bio, Work & Ideas

Connor Guerrero

Conference affiliation: Crusoe

On this page

Connor Guerrero is a machine-learning engineer and developer advocate whose work connects model development with the hardware and software needed to put models to use. His projects span low-power voice interfaces, accelerator-backed language-model applications, and distributed training that can recover from hardware failure.

From embedded speech to accelerated AI

At AONDevices, Guerrero worked in machine-learning and lead application engineering, developing low-power TensorFlow voice models and adapting speech systems to regional accents. That work brought together two practical demands: fitting useful model behavior into constrained devices and handling variation in how people speak.

His move into developer relations at Tenstorrent in 2025 shifted the emphasis toward helping other engineers build on AI accelerators. He worked on accelerator-backed vLLM reference applications and developer enablement, connecting model-serving infrastructure with applications developers could use as starting points. Rather than stopping at the capabilities of the hardware, this work addressed the next step: making those capabilities accessible through working software and understandable workflows.

Guerrero joined Crusoe’s developer advocates group in February 2026. His authored work there extends that practical focus to distributed GPU training, where keeping a workload running requires coordination between the model, scheduler, and infrastructure.

Making models actionable

Two of Guerrero’s public projects make the steps between model input and application behavior explicit:

  • Speech directly to structured commands: His End-to-End Spoken Language Understanding project is a fixed-command, resource-constrained proof of concept. A compact transformer maps speech directly to action, object, and location fields without first producing a transcript. Those fields describe what an application should do, what it should act on, and where. The limited command vocabulary distinguishes the project from a general voice assistant; training, evaluation, benchmarking, and an interactive interface give developers a way to examine the approach end to end.
  • Local function-calling agents: His Local AI Agent from Scratch notebook uses vLLM and an OpenAI-compatible API to expose prompt construction, response parsing, and function-call routing. It makes the transition from a language model’s response to an executable tool request concrete, giving developers an implementation they can inspect and adapt.

Recovery as part of the training workflow

Guerrero’s guide to self-healing distributed PyTorch training addresses a problem that infrastructure replacement alone cannot solve. A new GPU node restores capacity, but the interrupted training job still needs to be scheduled onto usable hardware and recover its saved progress.

The guide connects automated node replacement and checkpoint recovery using Kubernetes, Slurm, and PyTorch. Infrastructure remediation and job recovery form successive parts of the workflow: replace failed capacity, requeue the workload, and resume training from a checkpoint. This continues Guerrero’s progression from fitting voice models into constrained devices to making larger AI systems usable through concrete application and operational workflows.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Crusoe combines Slurm’s training scheduler with Kubernetes infrastructure management. The recovery path connects GPU failure detection, node replacement, job requeueing and application checkpoint loading—while researchers keep their familiar commands.

  • Slurm supplies training-aware scheduling and familiar job commands; Kubernetes and AutoClusters connect those jobs to infrastructure telemetry and automatic node replacement.
    1:28 ↗
  • Recovery crosses layers: Slurm cancels and requeues the job, AutoClusters replaces eligible failed nodes, and application code loads the checkpoint.
    10:03 ↗
  • The two-A100-node demonstration reports roughly five minutes for hardware replacement and less than fifteen minutes for end-to-end training recovery.
    11:35 ↗
  • Keeping GPU nodes in Kubernetes allows capacity released by completed Slurm jobs to serve inference workloads.
    15:01 ↗