Dylan Bristot is a technical product marketer and the creator of WhatLLM, a platform for comparing language models. His work at Nebius has included helping establish the go-to-market strategy for AI Studio and Token Factory. He connects AI infrastructure with practical developer decisions: which model to use, what a workload will cost, and how to improve a system after it reaches users.
From financial software to AI infrastructure
Bristot’s career spans financial software, cloud infrastructure, and AI products. At ActiveViam, he worked in marketing operations and product marketing, implementing Salesforce and marketing automation. In 2022, he co-founded Pfpedia, a Web3 platform addressing fragmented NFT-project discovery. He subsequently joined Scaleway, where his product-marketing responsibilities covered cloud compute and networking, including ARM-based instances and compute products intended for production workloads.
At Nebius, his responsibilities have included technical content, developer partnerships, and hackathons, combining explanations of the infrastructure with opportunities for developers to use it. In 2026, he led product marketing for Nebius Token Factory. In a joint presentation with developer advocate Sujee Maniyam, Bristot addressed the tradeoffs between closed APIs and self-hosting, and the cycle connecting inference, production data, post-training, and deployment. Maniyam covered the serving optimizations behind fast inference.
WhatLLM gives Bristot’s interest in model selection a concrete expression. He built the model-comparison platform to bring performance, pricing, speed, and latency into the same decision. His accompanying writing examines how those comparisons change when applications move from short conversations to agents that make repeated model calls and use tools.
Choosing models and improving production systems
Cost per completed task: Bristot argues that token prices alone poorly describe agent economics. Repeated context, reasoning, tool calls, and retries can make an apparently inexpensive model costly in practice. His approach to model selection weighs capability against task cost and the consequences of failure: demanding or consequential work may justify a frontier model, while repeatable workloads can use cheaper models with escalation. Some steps are better handled by a deterministic API call. He recommends testing shortlisted models on the actual workflow rather than choosing from one leaderboard column.
Production feedback as training data: Bristot treats deployment as the beginning of model improvement. In his writing on Data Lab, recurring prompts, failed outputs, and unusual user requests become material for the next training cycle. The process is practical: inspect inference logs, isolate relevant examples, curate a reusable dataset, and feed it into post-training. He emphasizes reducing the exports, scripts, and separate environments that slow iteration; connecting existing S3-compatible storage can also avoid creating another raw-data copy.
Agent reliability as system design: Bristot’s production-agent guidance connects model execution with orchestration, evaluation, and memory. Evaluation must catch misrouting, repeated mistakes, and dead ends; memory must retain useful context across interactions. Application safeguards include input validation, limits on runaway tool calls, and recovery from external API failures. These requirements explain why a convincing demo can still struggle with real users.
Live information alongside model reasoning: In writing co-authored with Tavily’s Jakki Jakaj, Bristot describes a research-agent pattern in which the model decides whether it needs current information, calls a search tool, and uses the returned material to form an answer. Search supplies fresh information and relevant context, extending model selection and serving into the surrounding components an application needs.
Dylan Bristot and Sujee Maniyam explain how Nebius Token Factory connects production data to model improvement, then walk through the hardware, routing, caching and decoding choices behind fast inference.
A production model needs a recurring path from inference logs through dataset preparation and post-training back to deployment.
Serving the same model on different hardware and engines can produce different performance and cost; engine selection belongs to the model-specific optimization work.
Cache offloading preserves reusable work outside GPU memory, while prefill/decode separation assigns different resource demands to separate GPU groups.