Tanay Varshney is a senior product research engineer at NVIDIA, working with NeMo and NVIDIA’s inference microservices to improve language models and agents. His work connects the practical demands of deploying AI—retrieving useful information, managing inference costs, and coordinating tools—with research into choosing the right model for each task.
From applied machine learning to agent systems
Varshney earned a master’s degree in computer science at New York University, focusing on computer vision, data visualization, and urban analytics. His early projects included subway-traffic analysis using MTA turnstile data and quadcopter localization using sensor fusion: applications that required turning imperfect observations into useful decisions.
At NVIDIA, his early technical writing addressed the steps between a trained model and a working service. In 2021, he co-authored a guide to adapting speech-recognition models through transfer learning, connecting domain-specific fine-tuning with deployment through NVIDIA Riva. The following year, he co-authored an end-to-end inference deployment guide linking TensorRT optimization with Triton model serving. The distinction mattered: accelerating a model still left developers with configuration, client interfaces, scaling, and infrastructure to handle.
By 2023, Varshney was developing a similarly concrete account of language-model agents. He framed them as systems combining a coordinating core, memory, tools, and planning. His agent architecture guide used financial analysis to make the need for those components tangible. Comparing revenue across reporting periods requires retrieving separate figures, keeping track of intermediate results, and calculating the difference. More interpretive questions require decomposing the task further and selecting the appropriate tools.
His subsequent work expanded into multimodal retrieval and reasoning systems. In 2026, he and Annie Surla received equal-contribution credit on research into routing with prefill activations. That research tackles a consequential question for AI applications: how can a system estimate which model will answer correctly before paying for the answer?
Making model capability useful
Activation-based model routing: Varshney and his collaborators use the internal representations a model produces while processing a prompt to predict the success of candidate models. Their encoder–target decoupling approach separates the open-weight model supplying those representations from the model being assessed, allowing the router to estimate the performance of closed-source models without accessing their internals. A shared neural predictor estimates correctness across the candidate pool; the routing decision combines those estimates with expected inference cost. This makes model selection sensitive to more than the apparent topic of a question.
Routing across an agent’s changing workload: Varshney co-authored the introduction to NVIDIA NeMo Switchyard, which separates routing decisions from provider endpoints. Its routing strategies can respond to the progress of a task: repeated errors or prolonged exploration can justify a more capable model, while steady implementation can favor a cheaper one. This approach treats model choice as an ongoing engineering decision within an agent’s work.
Multimodal retrieval with explicit tradeoffs: In his work on video and audio retrieval, Varshney favors converting information from different media into a common textual representation when that fits the application. Transcribed speech and visual descriptions can then share a retrieval pipeline. He also identifies the cost of that simplification: conversion can lose information, and actions unfolding across frames require different treatment from a tutorial whose individual slides carry the meaning.
Spending compute where it improves the answer: His co-authored work on retrieval reranking shows how an additional relevance-scoring step can reduce the amount of material sent to an expensive language model. His reasoning framework distinguishes longer reasoning, searching across candidate solutions, and iterative critique. Each has a different mechanism; generating more candidates also creates the separate problem of selecting a good answer.
Routing an agent means deciding who plans, who executes, what context they share, and when to switch—not merely choosing the cheapest model for a prompt.