Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI
Read the talk
Making Slower, Intelligent Models Work in Realtime Voice Agents
Bohan Li explains how EliseAI overlaps transcription, response generation, tool calls, and speech playback to shorten the wait a caller experiences.
From a talk by Bohan Li
At a glance
Ideas worth remembering
The cascaded stack creates three distinct opportunities to improve responsiveness: transcribe input sooner, prepare decisions earlier, and begin speech before the complete response is ready.
Speculation requires revision. New input cancels stale transcription corrections, and useful background tool results can cancel and restart an early response. The harness withholds speech until the caller’s utterance ends.
Cached speech prefixes hide waiting time while fresh synthesis handles personalized content. The provider still generates the full sentence for prosody, and the harness suppresses duplicate opening audio before playing the continuation.
The recorded booking exchange illustrates the intended experience, including waiting for the caller and changing appointment options. It does not establish backend booking persistence, numerical performance, or a guaranteed inaudible join between cached and newly generated speech.
A voice stack built around perception, planning, and control
Bohan Li introduces a voice-agent harness designed to deliver realtime conversation while retaining the intelligence needed from a capable language model. Drawing on his previous work in self-driving, he explains the choice of a cascaded architecture through three responsibilities: perception, planning, and control.
Perception converts signals from the world into data a reasoning system can use. In self-driving, Li associates this with cameras, lidar, and bounding boxes; in voice, it is transcription. Planning then takes that representation and decides what to produce. Here, the language model consumes the transcription and generates the agent’s response.
Control turns the planner’s output into an action in the world. A car converts a trajectory into driving controls; a voice agent converts text into audible speech. This separation gives Li three places to improve responsiveness. His stated goal is to accelerate each layer without sacrificing the intelligence of the resulting conversation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Fast transcription with context-aware corrections
The perception layer uses what Li calls a streaming speculative transcriber. A fast streaming engine supplies text promptly, while a slower, more accurate batch transcriber processes additional context and can correct that text. The design lets downstream work begin from an early interpretation rather than waiting for the slower engine on every update.
His example begins after the agent asks for a name and date of birth. The first short transcription arrives from the streaming layer. If the corrective layer produces the same text, it does not issue a correction. As the caller continues, newer streaming text can cancel an older corrective result: the older result may have been computed more carefully, but it no longer reflects all the available audio. Freshness therefore matters alongside accuracy.
The useful correction comes when the slower transcriber uses the question’s context to distinguish a name from a date of birth. That is a semantic improvement, rather than simply a change in formatting. Li treats subsequent punctuation-only detections as unimportant in this example, then releases the resulting text to the agent. He gives a qualitative account of improved recognition, without an accuracy measurement or a precise timing benchmark.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Generate early and move tool work into the background
At the planning layer, Li focuses on reducing sequential round trips through a slow but intelligent language model. Tool calling can require repeated inference, so background agents perform that work and insert its results into the main agent’s context. He describes the main agent as receiving context as though it had made the call itself. The mechanism is context insertion; he does not specify the message format used to represent that tool history.
Each transcription detection also triggers an early generation of the main agent’s response. Generation can therefore overlap the caller’s speech, but the harness withholds the response until it confirms that the caller has finished. This separates preparing an answer from speaking it. An early draft can be revised without interrupting the person who is still supplying information.
The name-and-date-of-birth example shows why those early generations remain provisional. An initial acknowledgment does not contain either field, so the background tool does not fire. A later detection still lacks a recognizable name. With more text, the main agent prepares to ask the caller to spell the name because it suspects a transcription error. That response is only a candidate: it has been prepared from incomplete information.
When the corrected transcription arrives, the background agent can finally find the name and date of birth. The harness cancels the generation made without those tool results and triggers a new one with the needed context. The tool work includes correcting name mistranscriptions and phonetic matching. Once the end of the utterance is detected, the response is emitted. This trades some discarded generation for a chance to have the useful response ready sooner; it does not eliminate inference or regeneration.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Start speaking with a cached prefix
The control layer receives text as a stream and needs to turn it into audio quickly. Li’s objective is to begin playback before the language model finishes generating the full response. Speaking the beginning while the rest is still being written hides some of the remaining generation time from the caller.
A prefix cache watches the agent’s text stream for word sequences whose audio has already been generated. Li says that audio can come from a prior generation or from generation within the same call. In the walkthrough, the cache does not act on the first word alone; it waits for more words and gets its first hit after three. At the same time, the text flows to Cartesia through a WebSocket connection. Cache matching and fresh synthesis proceed together.
The reusable opening in his example is a confirmation that the caller supplied a name. Adding the particular name causes a cache miss: the generic opening repeats across responses, while the personalized continuation may be new. At that point, the harness emits the cached opening as the remaining text continues to arrive. The cache has shortened the wait for speech to begin, even though fresh synthesis is still needed for the rest of the response. The example explains the opportunity for reuse but does not establish a cache-hit rate.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep sentence context while suppressing duplicate audio
The synthesis provider receives the entire response text supplied up to that point, including the opening already played from the cache. It does not know that a prefix cache exists and generates the full sentence with natural prosody. Retaining the opening gives the provider sentence context for the continuation, even though the caller will not hear that opening from the provider.
When the newly synthesized audio returns, the harness suppresses the portion corresponding to the cached speech and plays only the remainder. That continuation follows immediately after the cached audio finishes. Li acknowledges that there might be a small hiccup, while suggesting callers probably would not notice it. He does not explain how the cut point is aligned between the two audio versions or quantify the transition’s quality. The claimed seamlessness is therefore a qualitative result, rather than a demonstrated guarantee.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A booking conversation brings the harness together
Li then plays a recorded clinic-booking call. The caller says they might be pregnant and wants an ultrasound to confirm. The agent asks for a name and date of birth, establishes that the caller is a new patient, and requests permission to text a link for uploading insurance information. After the caller agrees, the agent says the link has been sent. These turns show the conversational sequence; they do not reveal the underlying tool requests or establish whether the recording uses a live clinical system.
The agent subsequently reports receiving the insurance information and offers Thursday, July 2nd at 10:00 a.m. The caller asks for a moment to check a calendar, and the agent agrees to wait. When the caller requests something for the following week, the agent offers Tuesday, July 7th at 2:00 p.m. or 3:00 p.m. The caller chooses 2:00 p.m., and the agent confirms that the appointment has been booked. The exchange illustrates a change of scheduling preference and a pause within the workflow, rather than only a fixed sequence of questions.
After the call, Li returns to the harness as the source of much of the conversational behavior. Streaming and background work support the simple exchange heard by the caller, and he emphasizes the engineering effort required to make it feel natural. The demonstration supplies a concrete workflow, but no numerical latency comparison or accuracy evaluation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The product goal: help in housing and healthcare
Li closes by placing the engineering work in EliseAI’s broader focus on housing and healthcare, which he describes as critical areas of people’s lives. He says the company is headquartered in New York and is working to expand its Bay Area presence. The final invitation to join the team connects the technical presentation to that practical mission before the talk ends with applause.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> My name is Bo. I'm going to be here
- 0:14
presenting real-time voice agents with
- 0:17
Frontier Intelligence. Effectively,
- 0:19
going to be talking a little bit about
- 0:21
how we at Xnor.ai architected our voice
- 0:24
agent harness to get real-time voice
- 0:27
with the Frontier level of intelligence
- 0:30
that we need.
- 0:32
Okay. So, before I start,
- 0:35
I think I wanted to kind of draw some
- 0:37
parallels about
- 0:39
why we decided to go with cascaded voice
- 0:42
agents and especially kind of
- 0:45
comparing that to self-driving cars
- 0:47
which I was working in before. So, to
- 0:50
me, cascaded voice agents makes a lot of
- 0:52
sense when you view it in lens of kind
- 0:54
of breaking it down into perception
- 0:56
which is
- 0:58
for self-driving cars, it's you know,
- 1:00
the bounding boxes, the camera, the
- 1:02
lidar.
- 1:03
For voice, it's going to be the
- 1:05
transcription. Basically, effectively
- 1:07
turning these like signals from the real
- 1:09
world into
- 1:11
elements of data that the language model
- 1:14
or whatever brain you're working on
- 1:16
can process.
- 1:19
Second one is the planning step which is
- 1:21
pretty straightforward.
- 1:23
This is where the language model
- 1:25
will take in the outputs from the
- 1:27
perception stage and produce the outputs
- 1:30
that you want to produce out back and
- 1:32
out into the real world.
- 1:34
And finally, there's the controls layer
- 1:36
where
- 1:38
on self-driving, you'd be taking the
- 1:39
trajectory that the planner would output
- 1:42
and kind of turn it into the real
- 1:45
controls to kind of build like drive the
- 1:47
car. Here, we're turning the text into
- 1:51
audio that we use to express our voice
- 1:53
agent's thoughts.
- 1:56
And um
- 1:58
yeah, so here I'm going to be like going
- 1:59
to diving into each one of these
- 2:01
elements and we've made a few kind of
- 2:04
interesting tricks on each of these
- 2:06
areas to
- 2:08
improve the speed of our voice agents
- 2:10
without sacrificing the intelligence.
- 2:13
So, the first one is going to be uh the
- 2:16
transcriber layer. So, we came up with
- 2:18
this concept called like the streaming
- 2:19
speculative transcriber where
- 2:21
effectively we are layering a fast
- 2:24
streaming transcriber like Flux on top
- 2:27
of or kind of below a
- 2:29
uh scribe V2 or a accurate batch
- 2:31
transcription which kind of takes in
- 2:34
more context. It's a little bit slower,
- 2:35
but it will give you more accurate
- 2:37
detections.
- 2:39
So, we're going to walk through a
- 2:40
scenario. So, in this in this case the
- 2:42
agent just asked, you know, providing
- 2:44
can you provide your name and date of
- 2:46
birth and the user is going to say this
- 2:48
and we'll see how that plays out um
- 2:51
timing-wise. So, first we're going to
- 2:53
get, you know, the short detection. Um
- 2:56
we'll get it from we'll get it from the
- 2:57
streaming layer. The accurate
- 3:00
layer uh the corrective layer is not
- 3:02
going to fire because it's the same
- 3:03
text.
- 3:04
Um we're going to get some more
- 3:07
streaming text detections and in this
- 3:09
case the corrective layer is actually
- 3:11
canceled because we got new um new text.
- 3:14
So, you know, more context, more audio
- 3:18
is going to beat the old accurate one.
- 3:21
And here's where kind of the first
- 3:22
correction comes in. So, because the
- 3:25
scribe V2 layer understands, you know,
- 3:28
the the context of the question, it's
- 3:30
able to understand that this is talking
- 3:31
about name and this is a date of birth.
- 3:34
Then a couple more detections, these are
- 3:35
just punctuation, we don't care.
- 3:37
And so, in the end we kind of release
- 3:39
this text over to the agent.
- 3:44
And moving on um to the language model
- 3:46
layer.
- 3:47
So, here since we're kind of using these
- 3:50
slow but intelligent LLMs, we really
- 3:53
want to reduce the number of round trips
- 3:55
and the thing that causes us to do a lot
- 3:58
of inferences is tool calling. So, one
- 4:00
way to get rid of that is by having
- 4:03
background agents do the tool calling
- 4:06
for you and kind of
- 4:08
um push the tools back into the context
- 4:11
of the main agent so that it thinks it
- 4:13
made the tool call, but
- 4:15
um
- 4:16
but it it it really didn't.
- 4:18
So,
- 4:19
uh so, we remember from like detections
- 4:21
from before.
- 4:22
So, well, what happened is each one of
- 4:24
these detections is going to trigger a
- 4:27
um an early
- 4:30
kind of generation of the agent and we
- 4:33
but we won't actually
- 4:35
emit this out until we're confirming
- 4:38
that the user has finished speaking. So,
- 4:40
in this case, the user says, "Sure." The
- 4:42
agent kind of knows that the user is
- 4:43
about to say something else. Our
- 4:45
background tool calling here, which is
- 4:47
going to be helping us find figure out
- 4:49
the name and the date of birth from the
- 4:50
user detection, is not firing. So,
- 4:53
nothing much there.
- 4:55
Um the next instant detection comes in.
- 4:58
It says that,
- 4:59
you know, still not really a name. Um
- 5:02
our agent kind of plays along and
- 5:03
continues there.
- 5:06
Now, kind of a more more context come
- 5:08
comes back. The agent kind of feels like
- 5:10
there should be a name. It's going to
- 5:12
ask to spell it out because it's
- 5:14
probably thinking there's some
- 5:15
transcription error here. Still no name
- 5:17
or date of birth.
- 5:19
And then finally, this you remember this
- 5:20
is kind of our corrected um final
- 5:22
instant detection from the transcriber
- 5:25
from the Scribe V2.
- 5:27
Um here, our eager kind of agent
- 5:30
generation that was made without any
- 5:32
tool calls is going to get canceled
- 5:34
because the background agent finally is
- 5:35
able to find the name and date of birth
- 5:37
it's looking for. So, it's going to
- 5:40
retrigger and now the the agent actually
- 5:42
has the context it needs.
- 5:44
Um and you see here, it's kind of we're
- 5:46
doing it the tool call here is a little
- 5:48
bit um um
- 5:49
some intelligence there. We're going to
- 5:51
be like, you know, correcting
- 5:52
mis-transcriptions of name, and doing
- 5:54
some like phonetic matching here.
- 5:57
Um and yeah, and then we'll kind of
- 5:59
once we've understood that this is the
- 6:02
end of the user utterance, we'll kind of
- 6:03
emit it out. So, pretty standard.
- 6:06
Okay, and then the next layer here is
- 6:08
going to be text-to-speech. So, with
- 6:10
text-to-speech
- 6:11
the goal is to kind of take what the
- 6:13
agent said, and the agent's going to be
- 6:15
emitting this in a streaming fashion.
- 6:17
So, we're going to need to
- 6:19
um produce audio as quickly as possible.
- 6:21
And ideally, what you can do is before
- 6:24
the agent has even finished generating
- 6:26
the full text you can have the audio
- 6:30
play, so it's kind of hiding the latency
- 6:32
of finishing the generation.
- 6:34
So,
- 6:36
um I'm going to kind of play the
- 6:37
streaming um
- 6:38
the stream the streaming uh agent output
- 6:40
now. So, starts with you.
- 6:43
And yeah, actually before I uh further,
- 6:46
there's this new concept that we're
- 6:47
introducing here called the prefix
- 6:48
cache. So, the prefix cache is going to
- 6:51
be looking at the
- 6:54
um agent stream, and seeing if we
- 6:57
already have generated audio for that
- 6:59
sequence of words um from like a prior
- 7:03
generation, or maybe like the same
- 7:04
generation
- 7:05
um in this
- 7:07
uh in this call as as well.
- 7:10
So, um it sees the word you. Uh we for
- 7:13
this prefix cache, we're going to be,
- 7:15
you know, we don't want to like
- 7:17
immediately hit on every single word.
- 7:19
We're going to be waiting for a little
- 7:20
bit more words.
- 7:22
Um so, after three words, the prefix
- 7:25
cache gets our first hit.
- 7:27
And um over here on the right, this is
- 7:30
kind of our text-to-speech standard
- 7:32
provider, you know, Cartesia is a
- 7:34
text-to-speech engine with web socket
- 7:35
support. So, we're we're piping the
- 7:38
agent through the the cache, and also
- 7:41
piping it through web socket.
- 7:44
Um more tokens come in, more cache, more
- 7:48
sending through web socket. Not much to
- 7:49
say here.
- 7:51
And okay, so now we get our first uh
- 7:55
kind of first unique thing, which is we
- 7:58
found the token that actually causes a
- 8:00
cache miss. And it makes sense. If we're
- 8:02
kind of caching previous generations,
- 8:05
um you said your name is is a pretty
- 8:06
common thing, but once you we add in the
- 8:08
name, suddenly we're that's that's going
- 8:10
to result in the cache miss.
- 8:12
At this point, we're actually going to
- 8:13
yield out our cached audio. So, you said
- 8:16
your name is is going to be
- 8:18
um emitted as the rest of the streaming
- 8:21
text is coming back. So, at this point,
- 8:23
the user hears the agent and user
- 8:25
doesn't really know what's going on.
- 8:27
They just looks like really fast
- 8:28
response times to them.
- 8:31
Um and now the kind of remaining text
- 8:33
flows through.
- 8:35
And at this point, we've already emitted
- 8:37
from the cache. The cache has done its
- 8:39
job. Um the rest we can kind of throw
- 8:42
into Cartesia.
- 8:43
And here's kind of the trick where
- 8:47
Cartesia has seen the entire transcript
- 8:51
up to this point. It
- 8:54
to to Cartesia, like it doesn't know
- 8:56
about the existence of this prefix
- 8:57
cache. It's just going to generate this
- 8:59
full sentence with, you know, standard
- 9:01
natural prosody.
- 9:03
But, what we do is when the generation
- 9:06
comes back, since we've already played
- 9:08
the audio here, we can actually suppress
- 9:09
the audio from Cartesia here and just
- 9:12
play out the remaining stuff. So,
- 9:15
the user, there's might be a tiny bit of
- 9:18
a hiccup. You know, I'll play some audio
- 9:20
later and you'll know that you probably
- 9:22
won't be able to notice.
- 9:24
And effectively, we're kind of taking
- 9:26
this and
- 9:28
or taking this part and just emitting it
- 9:31
directly after this audio has finished
- 9:33
playing. So, it looks seamless to the
- 9:35
user.
- 9:37
And yeah, so I guess at this point kind
- 9:40
of this is a lot of technical details.
- 9:42
I'll I mean the best way is just kind of
- 9:44
play a call and see
- 9:47
see this in action. So here I recorded a
- 9:50
call and let's just hear what happens.
- 9:55
>> This is Bo OBGYN Elise speaking.
- 10:00
>> Hi, I think I might be pregnant and I
- 10:02
want to schedule an ultrasound to
- 10:03
confirm.
- 10:06
>> That's exciting news. Can you provide
- 10:08
your name and date of birth?
- 10:12
>> Sure, it's Elise Trial.
- 10:14
Date of birth
- 10:15
2303.
- 10:21
>> Thank you. Are you a new patient with
- 10:23
us?
- 10:26
>> Yes.
- 10:29
>> Welcome to Bo OBGYN. Can I text you a
- 10:32
link to upload your insurance
- 10:33
information?
- 10:36
>> Sure.
- 10:39
>> Okay, sent. Let me know when you've
- 10:41
received it.
- 10:48
I've received your insurance
- 10:49
information. The earliest appointment I
- 10:52
have is Thursday, July 2nd at 10:00 a.m.
- 10:55
>> Actually, give me a second.
- 10:57
Let me check my calendar.
- 11:01
>> Sure, take your time.
- 11:04
>> Um do you have something for next week?
- 11:09
>> For next week, I can offer an ultrasound
- 11:11
with Dr. Avery Stone at North Clinic on
- 11:13
Tuesday, July 7th at 2:00 p.m. or 3:00
- 11:17
p.m. Do either of those work for you?
- 11:20
>> Yeah, 2:00 p.m. works.
- 11:23
>> Great. Your appointment has been booked.
- 11:26
We look forward to seeing you then.
- 11:29
>> Thanks. Bye-bye.
- 11:32
>> All right. Yeah, that's pretty much it.
- 11:34
Um
- 11:35
Yeah, you can kind of see our
- 11:38
all this like uh
- 11:40
streaming
- 11:42
and you know, a lot of things are
- 11:43
happening in the background and and you
- 11:44
know, this is what really makes like
- 11:46
voice agents interesting. And um there's
- 11:49
a lot of effort that can be done in the
- 11:51
harness to really kind of get a um
- 11:55
natural conversation, which is what
- 11:57
we're after.
- 11:59
Uh okay. Yeah, so I guess briefly, you
- 12:01
know, in the last part, I want to just
- 12:04
talk a little bit about Elise. So I
- 12:06
think Elise, you know, our headquarters
- 12:08
are in New York and kind of we're trying
- 12:10
to expand our presence here in the Bay
- 12:11
Area. Um we I think it's maybe like a
- 12:15
different style of company that I think
- 12:17
people are
- 12:19
like uh think of when they think about
- 12:21
AI startups in San Francisco. Where
- 12:23
we're actually very focused on um
- 12:26
just like helping people and helping
- 12:30
people where they need it, like kind of
- 12:32
the life's most critical areas. We work
- 12:33
on housing, health care and uh we're
- 12:37
doing really well and you know, here's
- 12:40
there's a link here to
- 12:42
um you kind of join our team and there's
- 12:44
going to we're going to be uh posting a
- 12:46
lot on Twitter, so you can follow us at
- 12:48
EliseAI as well.
- 12:50
Um yeah, that's that's it.
- 12:53
>> [applause]