← All speakers

Bio, Work & Ideas

Tanay Varshney

Conference affiliation: NVIDIA

On this page

Tanay Varshney is a senior product research engineer at NVIDIA, working with NeMo and NVIDIA’s inference microservices to improve language models and agents. His work connects the practical demands of deploying AI—retrieving useful information, managing inference costs, and coordinating tools—with research into choosing the right model for each task.

From applied machine learning to agent systems

Varshney earned a master’s degree in computer science at New York University, focusing on computer vision, data visualization, and urban analytics. His early projects included subway-traffic analysis using MTA turnstile data and quadcopter localization using sensor fusion: applications that required turning imperfect observations into useful decisions.

At NVIDIA, his early technical writing addressed the steps between a trained model and a working service. In 2021, he co-authored a guide to adapting speech-recognition models through transfer learning, connecting domain-specific fine-tuning with deployment through NVIDIA Riva. The following year, he co-authored an end-to-end inference deployment guide linking TensorRT optimization with Triton model serving. The distinction mattered: accelerating a model still left developers with configuration, client interfaces, scaling, and infrastructure to handle.

By 2023, Varshney was developing a similarly concrete account of language-model agents. He framed them as systems combining a coordinating core, memory, tools, and planning. His agent architecture guide used financial analysis to make the need for those components tangible. Comparing revenue across reporting periods requires retrieving separate figures, keeping track of intermediate results, and calculating the difference. More interpretive questions require decomposing the task further and selecting the appropriate tools.

His subsequent work expanded into multimodal retrieval and reasoning systems. In 2026, he and Annie Surla received equal-contribution credit on research into routing with prefill activations. That research tackles a consequential question for AI applications: how can a system estimate which model will answer correctly before paying for the answer?

Making model capability useful

  • Activation-based model routing: Varshney and his collaborators use the internal representations a model produces while processing a prompt to predict the success of candidate models. Their encoder–target decoupling approach separates the open-weight model supplying those representations from the model being assessed, allowing the router to estimate the performance of closed-source models without accessing their internals. A shared neural predictor estimates correctness across the candidate pool; the routing decision combines those estimates with expected inference cost. This makes model selection sensitive to more than the apparent topic of a question.
  • Routing across an agent’s changing workload: Varshney co-authored the introduction to NVIDIA NeMo Switchyard, which separates routing decisions from provider endpoints. Its routing strategies can respond to the progress of a task: repeated errors or prolonged exploration can justify a more capable model, while steady implementation can favor a cheaper one. This approach treats model choice as an ongoing engineering decision within an agent’s work.
  • Multimodal retrieval with explicit tradeoffs: In his work on video and audio retrieval, Varshney favors converting information from different media into a common textual representation when that fits the application. Transcribed speech and visual descriptions can then share a retrieval pipeline. He also identifies the cost of that simplification: conversion can lose information, and actions unfolding across frames require different treatment from a tutorial whose individual slides carry the meaning.
  • Spending compute where it improves the answer: His co-authored work on retrieval reranking shows how an additional relevance-scoring step can reduce the amount of material sent to an expensive language model. His reasoning framework distinguishes longer reasoning, searching across candidate solutions, and iterative critique. Each has a different mechanism; generating more candidates also creates the separate problem of selecting a good answer.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Routing an agent means deciding who plans, who executes, what context they share, and when to switch—not merely choosing the cheapest model for a prompt.

  • Which tasks justify the expensive model?
    0:21 ↗
  • Keep frontier judgment, delegate the implementation
    3:15 ↗
  • The task changes while the agent is working
    8:48 ↗
  • Cheap tokens can produce an expensive task
    13:53 ↗
  • A sidekick keeps its own running context
    18:23 ↗
  • Routing can happen inside a model artifact
    20:12 ↗
  • Remember where the evidence lives
    21:59 ↗
  • Cache-aware routing meets the always-running agent
    24:43 ↗
  • Local capacity and shorter context change the calculation
    29:53 ↗
  • Use supervision opportunities to detect trouble
    32:43 ↗
  • Measure confusion, then inspect the trace
    37:48 ↗
  • Learn from the routing decisions users correct
    41:16 ↗
  • What remains when models become better collaborators?
    43:04 ↗

References