← All speakers

Bio, Work & Ideas

Tarun Sunkaraneni

Conference affiliation: Amazon

On this page

Tarun Sunkaraneni works on the software surrounding machine-learning models: representing conversations for training, making model deployment practical, and keeping GPUs supplied with data. His multimodal training work, presented while affiliated with Amazon AGI at AI Engineer World’s Fair 2026, examines how image preparation and data movement can limit performance even when GPU capacity is available.

From dialogue research to developer infrastructure

Sunkaraneni graduated from the University of Utah with a computer-science bachelor’s degree in 2020, working in its natural-language processing group under Vivek Srikumar. His honors thesis, Transformer-Based Observers in Psychotherapy, explored how language models could classify patient and therapist utterances in motivational interviewing to support feedback on counseling practices.

The central challenge was representing dialogue for transformers. Models pretrained on written passages needed to distinguish speakers, conversational turns, and the particular utterance being classified. Sunkaraneni tested encodings that made those distinctions explicit, alongside training approaches that combined patient and therapist classification. The unified approach improved patient-label predictions in his experiments. The accompanying bert-therapy repository contains experimental code and runtime scripts.

The thesis also examined whether better predictions came with useful explanations. Gradient-based saliency maps gave disproportionate importance to special tokens, while attention visualizations did not consistently reveal interpretable patterns. These methods offered limited insight into why the models made their classifications.

During a summer 2020 internship on Plaid’s Domestic Integrations team, Sunkaraneni worked on an internal HTTP request library intended to reduce latency, improve parallelism, and support newer networking protocols. The work addressed infrastructure for financial integrations that had to accommodate changes in partner institutions’ systems.

He subsequently joined Microsoft. His involvement in Azure Bicep, the language for describing and deploying cloud resources, included an assignment to an issue concerning misleading errors caused by commas in arrays. The issue illustrates a practical developer-tooling concern: diagnostics need to help infrastructure authors identify the actual mistake in a configuration file.

Serving models and keeping GPUs fed

Sunkaraneni’s later projects address two different constraints on using models effectively:

  • Practical model serving: His vllm-openai-hf repository packages vLLM in a Docker image for Hugging Face Inference Endpoints. It connects the inference engine to a hosted deployment workflow through a custom container and configurable vLLM environment variables.
  • Data-pipeline performance: In his multimodal training example, a pipeline training a Qwen3-VL-style model on images stored in S3 initially spent about 85% of its time waiting for data. Fetching, decoding, and resizing images one at a time left the GPU idle. Concurrency allows preparation tasks to overlap; prefetching prepares data before the trainer requests it; and Ray’s object store avoids unnecessary copies of large image tensors between processes. The example raised GPU utilization from 15% to 90%.

His performance work also shows why an optimization must be judged at the scale where it will run. Once image preparation and transfers improved, a single machine’s network card became the limiting resource. Spreading workers across nodes and using zero-copy reads produced a further reported 50% increase in throughput. The useful lesson is to follow the bottleneck as it moves: faster model execution cannot compensate for a pipeline that fails to deliver the next batch.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Tarun Sunkaraneni walks through a multimodal training pipeline whose bottleneck moves from serial image loading to batch timing, tensor copies and finally a single node’s network card.

  • A batch that needs more than 100 seconds of fetching and only 15–20 seconds of training is primarily a data-pipeline optimization problem.
    5:38 ↗
  • Use concurrency that fits the work: asynchronous requests for I/O, threads where operations release the GIL, and processes for separate CPU execution.
    7:07 ↗
  • Prefetching hides preparation behind training when the producer stays ahead, while adding queue startup and checkpoint-resumption complexity.
    9:06 ↗
  • Passing object references lets training ranks retrieve large tensors directly, avoiding repeated copies through the coordinator.
    10:37 ↗
  • Evaluate configuration at the intended scale: spread scheduling and zero-copy retrieval gave no small-scale benefit here, but together added 50% throughput at large scale.
    14:15 ↗

References