AI Engineer World's Fair 2026
Latency Is a Budget. Humanlike Is the Goal. — Jesse Hall, LiveKit
Read the talk
Latency Is a Budget. Humanlike Is the Goal.
Jesse Hall explains why voice agents need conversation-level evaluation, audible latency measurements and asynchronous tools—and how to spend a response-time budget on behavior that feels human.
From a talk by Jesse Hall
At a glance
Ideas worth remembering
Evaluate the complete voice stack: a transcription error can drive a confident but incorrect downstream action.
Measure the wait until audible speech. First-byte and first-token metrics can hide leading silence and the time needed to finish a speakable sentence.
Separate slow tools from conversation progress so the agent can remain available while backend work continues.
Audio-based turn detection and provisional transcripts can overlap pipeline work. Hall’s practical response targets are within 1.5 seconds, with around 600 milliseconds feeling more human.
Conversation benchmarks should score correct outcomes, interruptions, tool behavior, privacy and unnecessary steps—not completion alone.
A caller interrupts, and the agent falls apart
A background agent can spend another ten seconds researching while its user does something else. A voice agent has someone waiting on the other end of the line. Jesse Hall, LiveKit’s staff developer advocate, opens with the familiar failure: a caller talks over an automated agent, and the conversation falls apart.
Shaving another hundred milliseconds off a model response will not teach the system to handle that interruption. Humanlike behavior requires listening while speaking, distinguishing a finished thought from a pause, and adapting when the caller changes their mind. Model selection and pipeline configuration have to serve those behaviors. Latency sets the available budget.
LiveKit supplies the real-time framework and infrastructure around these interactions, including audio, video and data. This talk concentrates on voice, where the system has to coordinate both what to say and when to say it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The best components can still book the wrong day
A conventional voice pipeline has three model slots: speech-to-text transcribes the caller, an LLM produces the response, and text-to-speech turns that response into audio. Integrated real-time models offer another option, though Hall notes that they are not suitable for every use case. Their availability adds another architecture to evaluate rather than settling the choice.
The problem with choosing each slot from a leaderboard is that errors travel through the pipeline. Suppose a caller requests an appointment on Tuesday the 3rd. Speech-to-text produces “Tuesday the 30th.” The LLM now receives the wrong date as its input and can confidently arrange the wrong appointment. The downstream stages do not automatically know that the transcription changed the request.
This is the “leaderboard trap”: a component score answers a narrower question than the one a deployed conversation presents. Hall contrasts clean benchmark audio with callers on speakerphones, in traffic, with children in the back seat. A model’s ranking cannot establish how the whole stack will handle those conditions. The useful selection question becomes: which stack should this agent ship?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Bytes and tokens arrive before the caller hears a response
Even after selecting the models, a dashboard can make the agent look faster than it sounds. The distinction is between the event a provider measures and the event that ends the caller’s wait. Two measurements need particular care:
- Text-to-speech: first byte versus first audible sound. An audio response can begin with silence. Hall reports measuring popular providers whose output contained hundreds of milliseconds of leading silence, reaching three-quarters of a second in the worst case. Receiving bytes during that interval does not give the caller a response.
- Language generation: first token versus first complete sentence. In the pipeline described here, speech synthesis starts with a complete sentence. A model can emit its first token quickly yet take longer to finish that sentence. Ranking models by token streaming speed can therefore produce a different choice from ranking them by readiness to speak.
Perceived latency follows the caller’s experience: how long until something audible happens? That pushes measurement across model boundaries. The LLM’s output must become usable speech input, and the TTS output must contain sound before either component’s speed helps the conversation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep the conversation moving while the booking system works
A useful agent also has work to do: look up an order, check availability or book a room. Those tools may take two seconds, ten seconds, or forty seconds or more. Faster language generation cannot make a forty-second backend operation disappear. A human receptionist handles this by acknowledging the work and remaining available while it runs.
The asynchronous tool pattern separates backend progress from conversation progress. In Hall’s explanation, the tool’s first update hands the microphone back to the agent while the operation continues in the background. The agent can keep speaking and listening rather than holding the conversation until the tool returns. As he puts it, “The conversation and the tool stop sharing a thread.”
The hotel demonstration makes the desired behavior concrete. The caller requests October 10–12 for two guests. While the agent is listing room types, the caller interrupts with a king-room choice. The agent moves on to the next relevant question—city or ocean view—rather than requiring the caller to wait through the menu or repeat the selection.
After collecting the ocean-view preference and booking name, the agent presents an October reservation priced at $582.40 and asks whether to proceed. The caller then changes the dates to December 20–24. The agent acknowledges the new range, replaces the quoted total with $1,164.80, and asks for confirmation again. Only after the caller says yes does it give its final confirmation. The observable change is a revised proposal before approval: the earlier dates and price no longer drive the next decision.
Parts of the demonstration were sped up for presentation, so its pacing cannot establish the agent’s unedited response times. Its useful evidence is the conversational behavior: accepting an interruption, retaining the room preference and handling a changed request. Those results depend on the framework, infrastructure and models working together.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Overlap the work between the last syllable and the first sound
Knowing when to speak introduces another tradeoff. An agent that starts too early cuts off a caller who is still thinking. An agent that waits for the complete transcript can respond too late. LiveKit’s audio-based turn detector uses speech cues such as pitch and intonation to identify the end of a turn without first waiting for all the text.
Hall reports a user-cutoff rate of 4.5% for LiveKit’s detector, compared with 9.9% for the best alternative at the same latency budget. This is a reported evaluation comparison, rather than a guaranteed rate for every deployment; the practical next step he recommends is running the open-source evaluation harness on your own data.
The response path can also overlap work. End-of-turn detection fires from audio before the final transcript arrives. The LLM is already working with a provisional transcript; if the final words match, Hall describes a saving of half a second. Speech synthesis then starts when the first complete sentence is ready, without waiting for the entire response. The provisional-text shortcut depends on that match—the recording does not describe the recovery path when the final words differ.
Where does the overlap save time? The diagram follows audio and transcription along separate paths that meet at response generation. The important relationship is that the final transcript need not hold up all language-model work, and the full language-model response need not hold up the first spoken sentence.
The window that matters runs from the caller’s last syllable to the agent’s first audible sound. Hall gives a practical target: within a second and a half end to end, the conversation can still hold, though it is near the edge; closer to 600 milliseconds starts to feel human. The design question is how much conversational quality fits inside that window. Once the response is fast enough, the remaining budget can support better model behavior.
The caller reaches the last syllable of the turn.
Audio-based turn detection and provisional transcription allow work to overlap. The caller’s wait ends at audible speech.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Score reservations, interruptions and unnecessary work
A conversation-level benchmark starts with a complete working agent. LiveKit’s first example is the hotel receptionist, backed by a real booking system. Hundreds of simulated callers interrupt, mumble and change their minds. These scenarios exercise the coordination that separate STT, LLM and TTS scores miss.
The scoring asks what a competent human receptionist should accomplish:
- Correct outcome. Did the intended reservation reach the database, with the correct name and dates?
- Conversational behavior. Did the system cut off the caller?
- Tool and privacy behavior. Did it invent tool calls or leak personal data?
- Efficient completion. Did it ask unnecessary questions or take ten steps for work that should have taken two?
Completion alone is too easy to game. An agent can guess its way to a completed task while making the interaction worse. Penalizing unnecessary questions and excessive steps makes the benchmark sensitive to how the result was reached, as well as whether a result exists.
To compare stacks, the benchmark keeps the agent and scenarios fixed, swaps the models, and supplies identical inputs byte for byte. That makes differences easier to attribute to the stack under test. It still measures performance on that particular agent: a hotel result does not automatically select the right models for every other workflow.
The methodology, evaluation sets and example agent are described as open source, so teams can apply the approach to their own agents. The central change is ownership of the test: evaluate the conversation your users will actually have, rather than borrowing a component ranking as a deployment decision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Spend the remaining budget on quality, then keep testing
Hall positions LiveKit as a framework that accepts different model providers, then previews a voice-tuned model on LiveKit Inference, identified in the supplied transcript as “Gemma 4:31B”. The tuning targets long prompts, many tools and tight response-time budgets—the conditions of a production voice agent rather than a short text exchange.
For that preview, Hall reports an 88% voice-evaluation pass rate and speech starting in about 380 milliseconds, more than twice as fast as the next model in the presented comparison. Those figures describe the announced evaluation, whose full comparison conditions are not developed here; they do not establish a universal ranking. The model was presented as launching the following day.
The ending turns the latency budget into an implementation checklist:
- Set the budget first. Decide how long the caller can wait for an audible response.
- Measure perceived latency. Include the work required to reach usable speech, rather than stopping at the first token or byte.
- Buy quality with the remaining time. Choose better model behavior when additional speed no longer improves the interaction.
- Run slow work asynchronously. Hall recommends making anything over a second asynchronous so the conversation can continue.
- Test whole conversations continuously. Evaluate the agent’s behavior as a system, and keep testing in production.
The final test is ordinary human behavior: a caller talks over the agent. A successful real-time system keeps listening, adapts and continues the task. Building toward that moment gives the budget, metrics and model choices a concrete purpose.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Extends the evaluation discussion into improvement: execution traces and task-specific feedback guide changes to prompts and agent programs, while human annotations can shape the evaluator.
Read the complete timestamped transcript
- 0:12
All right, let's get started. So most of the agents that we build today, they do their work in the background, right? They write code for us, they do research. Uh, you give them a task, you walk away, you come back. And so nobody really cares if, uh, a background agent takes an extra ten seconds or even ten minutes, right? We're used to that. But in-- at some point in this past year, you've probably talked to another kind of
- 0:42
agent, a real-time agent. Um, maybe, uh, it was a pharmacy refill, maybe at an airline. Um, I talked to one the other day when I called a car dealership, and, and so something happened on that call. You probably did the most human thing possible. You talked over that AI, and that voice agent fell apart and it just didn't feel natural, did it? And so, um, if you ask, uh, anybody in this industry, you're
- 1:12
probably gonna get a very similar answer of how to fix that. It's gonna be speed, uh, faster models, lower latency, shave off another hundred milliseconds. Um, but that is the wrong answer. Uh, and so latency, it is a budget, and I'll, I'll tell you exactly what that budget is, but the goal is something else entirely. The goal is h- being-- the voice AI being as human-like as possible. Uh, an agent that listens while it talks, uh, one that knows the difference between
- 1:42
you finishing your thought and you still thinking, uh, one that rolls with it when you change your mind. And so speed alone can get you none of that. What gets you there is which models you pick and how you configure your pipeline. And so that's the first problem that we're going to encounter. No one tells you how to pick the right models. The leaderboards can't help you. The vendors definitely can't help you. So let's talk about how we're going to fix that issue.
- 2:12
So a show of hands quickly, who here has ever built a voice agent and shipped it to real users? Okay. All right. Some of you have experience and you know the challenges that you'll face when deploying a real-time agent. Another show of hands, who here has used ChatGPT's advanced voice mode to talk with ChatGPT? Yeah, many of us have, have used that, right? So OpenAI actually chose LiveKit to power that.
- 2:42
LiveKit is the open source real-time framework and infrastructure for the most human-like voice agents possible, uh, agents that can hear, see, and speak with the person on the other end. Uh, and so not just voice, but also video, uh, data, and even physical AI. But because voice is, is challenging, uh, we're going to actually focus on voice during this talk. And so, um, what we're going to, uh, look at
- 3:12
is, is voice agents. They basically have, um, three models, uh, that work together as one system. So your speech-to-text, your LLM, and then your speech-- uh, text-to-speech. So three slots. Now, I do have to acknowledge that there are real-time models as well and, and real-time models, uh, they are a newer concept. They're getting better and better, uh, but they're not right for every use case. And just the fact that there are real-time models even adds more complexity
- 3:42
to our problem of trying to figure out which model is best for our voice pipeline. So here is what you have to choose from, hundreds of models, dozens of providers. A-and show of hands here, who has picked a model just because they saw it at the top of a leaderboard? Yeah, most of us. Um, but here's why that doesn't work for real-time agents. Leaderboards score components, but you are shipping a cascade, a
- 4:12
pipeline. So let's say that your caller wants to book an appointment on Tuesday the 3rd. Your speech-to-text hears Tuesday the 30th, and so from that moment on, it doesn't matter how smart your language model is, it's now confidently booking the wrong day and every error flows downhill from there and the layers below it have no idea. And so those leaderboard scores, they came from clean audio in a quiet room,
- 4:42
uh, but your callers, they're calling on a speakerphone in traffic with kids in the back seat. So those leaderboards never met your users. Uh, and there's a name for this, the leaderboard trap. So this-- the score measures the model, but again, you're shipping a conversation. But everyone building voice agents, they're asking, you know, "What is the best speech-to-text model, LLM, TTS?" But the real question is, "Which stack
- 5:12
should I ship?" So the-- these are two completely different questions, and the only-- uh, only one of them has a leaderboard, until today. Uh, we're gonna talk about m- that more in a bit. So let's say you fall into this trap. You pick your models based on the leaderboard and now you start measuring things. This is the next problem. You're probably measuring the wrong thing. You see, your TTS provider, uh, reports time to first byte, and that looks great on the dashboard, uh,
- 5:42
but bytes are not sound. And so we've measured popular providers that are actually shipping hundreds of milliseconds of silence at the beginning of their audio and, in the worst case, up to three-quarters of a second of nothing but silence before that first audio comes in. So time to first audio is actually what your caller hears. And so this is the distinction that this entire talk hangs on, this, uh, measured latency versus perceived
- 6:12
latency. So what your dashboard says versus what your user hears. The dashboard says it's really fast, but your ears say it's slow. Same story for language models. Time to first token is a text metric. Uh, so your TTS can't start speaking though on a token, right? So it needs an entire sentence. So that, that number matters for, for, uh, voice because, uh, it needs the, the entire first sentence. But the
- 6:42
model ranks differently on that. So a model that streams fac- fa- a model that streams tokens the fastest is not necessarily the model that finishes a sentence first. Now, if, if you-- let's say you do measure all of this perfectly, uh, there's another part of the conversation that, uh, the benchmarking that gets completely ignored and that is that a real, uh, agent doesn't just chat. It needs to do things, right? It has to look up orders, check
- 7:12
availability, book rooms, right? And so a tool is going to take time, two seconds, ten, forty or more seconds, and no model choice is going to be able to hide a s- a forty-or-more-second, uh, latency. So think about what a human receptionist would do here. They, they don't freeze while the, the booking system loads, right? They say, "Give me a sec. I'm looking up your, your account," right? They, they buy time and, and they listen while they
- 7:42
keep working, and that is that human-like behavior that we're looking for. So why can't an agent just keep talking and do things asynchronously? Well, it can. And here's the shape of that, and this is the only code that you're going to see. Uh, but this is a normal tool. Uh, there's just one thing different, and that is this line right here. Uh, the first update hands the mic back to the agent while it keeps, uh, working in the background. And that's the whole trick. The
- 8:12
conversation and the tool stop sharing a thread. Uh, and so don't take my word for it. Y- listen for yourselves.
- 8:21
Thanks for calling the LiveKit Hotel. How can I help you?
- 8:24
Can I book a room?
- 8:26
Of course. What dates are you looking at?
- 8:29
October the tenth through twelfth.
- 8:32
October tenth to twelfth. And for how many guests?
- 8:35
Two.
- 8:38
I have queen with two beds, king, double queen-
- 8:41
A king room, please
- 8:42
... a suite, or-- Would you prefer a city or ocean view?
- 8:45
Ocean.
- 8:48
Who should I put the booking under?
- 8:49
Jesse Hall.
- 8:55
You're all set for October tenth to twelfth in a king room with an ocean view for five hundred eighty-two dollars and forty cents on the card ending in one one one one. Shall I go ahead?
- 9:06
Actually, you know what? I need to change the date to December the twentieth through the twenty-fourth.
- 9:12
December twentieth to twenty-fourth. That brings the total to one thousand one hundred sixty-four dollars and eighty cents for December twentieth to twenty-fourth. Shall I go ahead?
- 9:23
Yes.
- 9:25
You're all set.
- 9:27
Okay. So we had to speed up some of the parts there because of time, but the, the con-- the agent kept the conversation going. There were no awkward pauses. The caller even interrupted it, and it kept going. Uh, and the caller also changed his mind in the middle of things, and the agent just rolled with it, right? This is the human-like behavior that we're looking for. So this is, again, not one model being good. This is an entire system being good. This is the infrastructure, the framework, and the
- 9:57
models all working together, and that's exactly what no, uh, component or model benchmark can see. Now, there's one more axis to this that the leaderboards miss entirely, and this is knowing when to speak and when not to speak. So humans negotiate turn-taking in a matter of just a, a f- a couple of hundred milliseconds, mostly from tone. And a pipeline that is-- waits for the entire transcript to come back, it's already late, right?
- 10:27
So we built a turn detector that is based on audio itself. So it listens to the rhyme, the pitch, the intonation of the speech of the audio, and then, uh, we published the numbers on this. So at the same latency budget, the best alternative cuts users off nine point nine percent of the time. Ours is four point five percent. And the eval harness is open source. Uh, you don't have to trust me. You can, uh, run your own data through
- 10:57
this and see what you get. So let's go one level deeper for a minute. The gap between your caller's last syllable and your agent's first audible byte. So end-of-turn detection fires from audio before the transcript lands. The language model is already running on a provisional, uh, transcript, and if the final words match, then you just saved a half a second. And then the TTS starts on the first complete sentence, not the last
- 11:27
token. And so, uh, the, the, the first audible byte is when the human stops waiting. And so every model that you choose, it shows up here in this window, and this window has a fixed size. If you can stay within, um, a second and a half end to end, the conversation still holds, but you're, you're right on the edge. If you can get closer to six hundred milliseconds, then it starts to feel really human. And so the question, uh, is never how fast can I
- 11:57
go, but what is the most human-like agent that I can fit inside this window? So back to latency as a budget, but human-like is the goal. So if we add all of this up, the question that we asked at the beginning, what models should I pick? That was never actually answerable, and so we built the thing that we wish existed. Here's how we built the benchmark. We take a complete working agent, the first one is that ho- uh, hotel
- 12:27
receptionist that you just heard, and it, it, it takes real bookings, it has a real back end, and it's also open source. You can look at the repo. And we run hundreds of simulated callers against it. Um, the simul- the simulator interrupts, it mumbles, it changes its mind just like a real caller would. And then we don't score it on vibes, uh, but we score outcomes, uh, defined the way that we would define them for a human. And so did the right
- 12:57
reservation land in the database? Did a caller get cut off? Uh, are the names and dates correct? Uh, are there-- were there any invented tool calls, uh, or any leaked personal data, right? So we measure all of those things against human equivalent behaviors. That's the bar. And then completion alone isn't enough because completion can actually, uh, be gameable, right? So a, a model can kinda guess its way to complete, right? And so we penalize what a completion
- 13:27
c- a completion score can't see. Um, if-- like, if it asks questions that didn't need to be asked, or if it takes ten steps to do something that should have only taken two steps. And then we swap the stack out. The same agent, the same scenarios, but different models. And every model sees the exact identical input byte for byte. Um, but this creates benchmarks against our agent example. Um, there are as many
- 13:57
benchmarks as there are agents, so the benchmark is actually yours. You can run this exact thing, uh, run your agent through this exact benchmark and get your scores. And so this is how you catch surprises. The whole thing is open source, the methodology, the eval sets, the agent itself. Uh, you can point the same thing at your agent and let us know how it performs for you. Now, here's an important fact about LiveKit, is that we don't make STT, LLM, or TTS
- 14:27
models. Every provider plugs into our stack, and so we don't have a horse in this race. We just watch all of them run a lot, um, and so we can confidently pick the winners. And so what you're looking at right here is a sneak peek of, of the Gemma 4:31B model on LiveKit Inference. Uh, but this isn't just an off-the-shelf Gemma 4:31B. It's tuned specifically for production voice agents with long prompts,
- 14:57
many tools, and tight latency budgets. It passes voice evals at eighty-eight percent, and it starts speaking in about three hundred and eighty milliseconds. So that's over two times faster than the next model, um, making this the act- the fastest, um, high-quality voice LLM available today. And so this is actually gonna- going to launch tomorrow on LiveKit Inference. So again, our goal is to have the most human-like agents possible while
- 15:27
sticking to that latency budget. So if you remember one thing from this talk, make it this one. A voice agent isn't a chatbot with a speaker s- uh, stacked onto it. It's human-like infrastructure, a real-time system that senses and responds the way that a human, a person would. The moment that you treat it that way, every decision is going to get clearer, the budget, the metrics, the models, and the bar that you test against. So here's the checklist. You can take a photo of this.
- 15:57
Set your latency budget first. Measure perceived latency. Spend whatever budget is left on model quality, not on speed that doesn't matter. And then anything over a second, it should go asynchronously so that the conversation never stops. And then, uh, just test the conversation, not the components, and then keep testing in production. So here's the benchmark. Uh, scan the QR code. You can check out the benchmark. You can run this against your own agents
- 16:27
and at, uh, remember, at some point in the near future, you are going to call a voice agent, and you're going to talk over it, and it's not going to fall apart. It's going to act like a human would in that scenario. So go build that. Thank you.