AI Engineer World's Fair 2024
How to build the world's fastest voice bot
About this talk
Daily CEO Kwindla Hultman Kramer explains how to build low-latency conversational voice agents by combining WebRTC audio transport, fast transcription, phrase endpointing, colocated LLM inference, and voice synthesis. Healthcare-intake and interactive geography demonstrations illustrate production voice-agent behavior and tool calling, and he introduces Pipecat as an open-source orchestration framework.
Chapters
- 0:00Daily, GPT-4, and early conversational voice bots
- 1:41Production voice-AI infrastructure and endpointing
- 5:25Healthcare intake demonstration and production tool calling
- 10:42Colocating transcription, inference, endpointing, and speech output
- 14:54Interactive geography voice-agent demonstration
- 18:36Pipecat and the future of voice-AI orchestration
Talk transcript
- 0:00
[upbeat music] So, uh, my name's Kwindla.
- 0:15
I, uh, work at a company called Daily. We make real-time audio and video infrastructure. And when GPT-4 was released last year, we started building lots and lots of experiments.
- 0:25
Uh, we thought GPT-4 was so good at conversational exchange in text that it really should be able to talk to you, literally talk to you, with your voice. So we built a bunch of things like this.
- 0:40
This is probably a year ago, and in Gen AI terms, it's ancient, ancient history, but probably worth looking at as context.
- 0:47
I'm your friendly chatbot, here to assist you with any questions or tasks you have in mind. How can I help you today?
- 0:54
Uh, can you tell me a joke?
- 0:57
Absolutely. Why don't skeletons fight each other?
- 0:59
Oh, you always start with that one. Give me a different joke.
- 1:03
All right, I'll give it another go. What do you call fake spaghetti? An imp-
- 1:07
Oh, I've heard that one too. Come on, maybe just- [laughs]
- 1:10
So we use tell me a joke from, you know, very early on all the way through today, and I think we know all the jokes that all the LLMs we use regularly, uh, tell, uh, which is funny, uh, but not necessarily funny in the same way they were the first time we heard them.
- 1:24
Um, so this is sort of a really high-level schematic of what we're trying to do here, right? We've got a user on a phone or a laptop. They wanna talk to their device, and then somewhere in the cloud, we've got a bunch of GPUs, uh, that are doing a whole lot of, uh, heavy duty compute, and we
- 1:41
need to talk to those, uh, that cloud computing resource somehow.
- 1:47
Uh, so as soon as you build stuff like the video we just saw, a couple of things come very much top of mind. One is speed really matters, and the other is architectural flexibility is really, really important.
- 2:00
I'll talk about both of those things today. Uh, let's start with architectural flexibility. So that really nice clean diagram gets really messy really fast. This is not, uh, unusual in an engineering, mm, software development problem domain, uh, but I thought I'd kinda make a slide based on, like, I looked at a bunch of source code.
- 2:20
I thought about all the conversations I've had with, you know, my colleagues and our customers and friends who are building this stuff. And it turns out you have to, at some level, kinda be aware of a bunch of these things if you want to build real-time, robust voice AI stuff, deploy it, scale it to production.
- 2:38
Um, it's a little bit of an intimidating map. We're definitely putting the multi in multimodal AI here, all the way from audio processing, things like echo cancellation, uh, and CPU management when you're encoding and decoding audio and video through networking issues like firewall traversal, all the way through to building things like, uh, retrieval augmented generation and tool
- 3:00
calling so that your real-world applications are really, really useful.
- 3:04
We can collect that kinda messy map into a few, a little bit more kind of, uh, uh, high-level categories. Um, it's worth going through these just really quickly because I think they give you a sense of what that map is.
- 3:15
So you need really robust and low latency media processing and transport. You've gotta encode the media. You've gotta send it over the network. That's gotta work really well. It's gotta work really fast.
- 3:25
Um, you need really good and fast transcription, uh, at least until the future of truly multimodal, uh, audio native models comes, which will happen at some point. Uh, and even after that, you probably are going to need to go from audio to text for lots of kinds of AI use cases.
- 3:43
Um, you have to do lots of real-time data pipeline and buffer management. So, uh, I think in Discord, I've probably maybe twenty or thirty times answered the question, "Why is my audio stream not working?"
- 3:55
uh, when I did local development on my Mac, and then I pushed it to an Intel box in the cloud. And it's 'cause, you know, Indian issues always get you if you're writing low-level audio stuff.
- 4:05
But there's a lot of sorta pipeline and buffer management that comes into play here. You wanna be able to swap between models for a whole bunch of reasons or use multiple models together.
- 4:15
Uh, you generally need to call out to external systems. You have to do things like phrase endpointing, which is the fancy academic term for when is the person done talking and when do they expect the AI to talk.
- 4:26
Uh, you need to handle interruptions really gracefully, and that actually is a whole rabbit hole of its own. Uh, people will interrupt the bot, uh, and you need to figure out what you're gonna do and how you're gonna handle that and how you're gonna maintain state when people interrupt.
- 4:40
Uh, you have to do echo cancellation unless you can convince everybody in the world who's using your thing to wear headphones a hundred percent of the time. Uh, and I've been doing audio-video development on the internet for a really long time.
- 4:50
Like, our core product at Daily is a WebRTC tech stack, so I can tell you that you will not convince everybody in the world to wear headphones all the time.
- 4:57
Uh, maybe if you're only deploying to people in, like, a call center professional context, but even then, it's pretty tough. Uh, and you need good, fast text-to-speech or voice generation on the other end.
- 5:10
W- The good thing is it's a lot, but when it all comes together, it really is amazing to build with these new gen AI tools. You can build stuff that feels magical and new to people, but that is genuinely useful, that sort of just feels like how things should be.
- 5:25
Uh, we built-- This is an example of something we built for a healthcare customer that solved a real pain point. I think if you've been to the doctor, which I'm sure all of us have, you've had to fill out way more forms than you thought you should at subsequent doctor visits.
- 5:38
And maybe you filled those, uh, out on paper. Maybe you, you know, got a link to a form to fill out before, uh, the visit, which is, like, the current state-of-the-art.
- 5:46
What if an LLM could just call you on your phone or you could click a link in an app and you can have a conversation and you never had to fill out a form again? [upbeat music]
- 5:56
Hello, Chat. My name is Jessica, and I am with Tri-County Health Services. To ensure your privacy and to confirm your identity, could you please provide me with your birthday, including the year?
- 6:07
Uh, yes, [REDACTED:dob].
- 6:11
I appreciate that, Chad. Can you also tell me-
- 6:13
Oh, sorry. Yeah, uh, [REDACTED:dob]
- 6:19
Thank you for providing your birthday, Chad. Now, could you please list your current prescriptions, including the medication name and the dosage for each?
- 6:27
Uh, yeah, I take lisinopril and also gabapentin.
- 6:33
Great. Could you please provide the dosage-
- 6:35
Oh, yeah, uh, sorry, ten milligrams and, uh, a hundred milligrams respectively.
- 6:43
Thank you for that information, Chad. Additionally, do you have any allergies that we should be aware of? If so, please list them.
- 6:50
Uh, yeah, I'm allergic to latex, and I'm also allergic to penicillin.
- 7:01
Thank you for letting us know about your allergies, Chad. Next, Har-
- 7:05
So I think you can see how all those pieces we just talked about are actually necessary to pull together something like that. But when you do, it feels simple.
- 7:13
It feels like it's supposed to. Um, that's a real video. Nothing in there is edited. Uh, the call-outs to the external system are mocked. They're not-- There's not actually a real electronic health record systems on the back end, but it's, it's, it...
- 7:26
They're, they're real tool calling. Um, and that stuff is now starting to move into production, uh, in, in the real world. So I think in a lot of cases, the GenAI frontier is moving very quickly, and people are really sort of pulling it, uh, out of our hands, which is super exciting.
- 7:42
Um, so let's switch gears a little bit and talk about fast because that's the other really big piece here. How fast do we need? How fast is fast? A-and also just a quick side note.
- 7:52
So we're talking-- There are always two kinds of fast when you're talking about engineering things. There's throughput, and there's latency. These days, for conversational interactions, throughput is pretty okay for all the tools we all use today.
- 8:05
Like, L-LLMs and other tools can generate content as fast as people can read it
- 8:11
or listen to it. But what's hard is latency, and latency is that sort of time to first byte, time to first token. Uh, in lots of, lots of engineering contexts, there's trade-offs between throughput and latency, complicated relationships between throughput and latency.
- 8:24
Uh, one of the graphs that I sometimes show in these talks is that, uh, throughput tends to improve by an order of magnitude every couple of years in lots of domains.
- 8:33
Latency improvements tend to be linear [chuckles] and, like, way behind throughput improvements. So latency's hard, and latency's mostly what bites us here. Human conversational latency, like if I am talking to another person, it feels weird to me if that person doesn't respond in about half a second.
- 8:51
Sometimes people respond actually a lot faster. We seem, as humans, to be doing, like, speculative decoding, next token prediction, just, like, natively. Like, that's what we do. I know what you're gonna say four or five words before you finish saying it.
- 9:05
I'm queuing up my response. I'm sort of doing my inference in this, like, greedy fashion. If you say something I didn't expect, well, I can, like, reroute. But most of the time, I'm right, and most of the time, if you actually record people in conversation, they'll respond in, like, two or three hundred milliseconds commonly.
- 9:21
And if they don't, they'll give you some kind of cue. So the, the sort of five hundred millisecond target is, is pretty important because we hit that uncanny valley pretty quickly when we're above it.
- 9:31
In fact, I think that video o-of my colleague, Chad, that you just saw, if you watched it with a critical eye, what I hope you saw was pretty cool orchestration of, like, state-of-the-art GenAI stuff and probably slower response times than really should be there.
- 9:49
Uh, so we spent the last couple of months-- I've spent the last couple of months really sort of thinking a lot about how to improve these response times. And just as a kind of benchmark or, like, relative, uh, like another number that shows how hard this is, like Gemini Pro's time to first token, it's like nine hundred
- 10:05
milliseconds. So if you're aiming for five hundred milliseconds, y-you're already almost double even before you do anything else, even before you send stuff, you know, o-over the network for other services or anything.
- 10:17
So what models and tools you choose are constrained. They matter a lot. It matters a lot how you string them together. Um, so just to pop up a level, again, this is what we're trying to achieve.
- 10:29
And the most powerful tool we have today for making everything run fast in this domain is actually putting as much together into one compute container as we possibly can.
- 10:42
So if the, if the really, really big things we're trying to do are natural language, uh, speech-to-text, and then phrase endpointing, so when should the bot do processing or talk, and then LLM inference, and then voice output, if we can put all those things together and run them locally and co-located, we're, like, way ahead o-of where we
- 11:04
are if we can't do that. And this is worth emphasizing because I, I think, m-I'm sure, like ninety-five, ninety-eight, ninety-nine percent of stuff we're all building today with GenAI, we're calling out to hosted services.
- 11:17
There's a lot of really, really good reasons for that. Uh, but that's tough in this domain if latency is what you're prioritizing. And latency might not be what you're prioritizing, and that's okay.
- 11:28
Like, there's lots of different trade-offs you can make. But if you're trying to make things really, really, really fast, you need to figure out how to host stuff yourself and how to host stuff in a way where you can tune and control and combine everything.
- 11:39
Um, so this is the part of the talk where I, like, look at the clock, and I look out at all of you, and I try to figure out how much tolerance you have for me talking about latency [chuckles] because I will, maybe ironically, will talk about latency for hours and hours.
- 11:53
It's what I'm obsessed with as an engineer. Uh, I do think it's worth just quickly ki-kind of going over this list of, like, kind of the best you can hope for latency numbers for a typical voice AI context, uh, 'cause some of them are non-obvious.
- 12:08
So first, what are we actually measuring? We're measuring the time-- Like, what do we really, really care about? We measure the time I stop talking. So if there's, like, a green waveform on one side and a purple waveform on the other side of this, like, you know, uh, audio, uh, editor, um- The time I stop talking, and
- 12:27
then there's some kind of gap, usually silence. We could play hold music if it's too long a gap. Um, and then there's another waveform on the other side when I first start to hear the LLM talking to me.
- 12:37
That's the gap we care about, the voice-to-voice latency, and that has to include everything. It has to include audio encoding, sending stuff over the network, all the processing, sending stuff back, playing it out the speakers.
- 12:49
The, the very first number here is actually kinda shockingly high. If you're using the laptop mic on macOS, it's like thirty, forty milliseconds of delay. There are some good reasons for that.
- 12:59
There's a lot of audio processing, again, echo cancellation, background noise removal, uh, gain control, all that stuff plays in, plays a role here. Uh, as these kind of use cases get more and more important, maybe hardware and operating system manufacturers will push this stuff down, but today you're sort of forty milliseconds into your five hundred millisecond target
- 13:18
time before the bytes ever even get into any of what we think of as software we control. Um, and then you're encoding the audio, and you're sending it over the network.
- 13:28
You're doing some packet handling. You're reassembling those packets on the other side in what we call a jitter buffer. You're decoding, and then you're finally at the, like, transcription stage.
- 13:40
If you are really careful about transcription and end pointing, you can probably get that down to a couple hundred milliseconds. Then you're doing LLM inference. In this context, you're really hoping you can get your time to first token from the LLM in maybe another hundred milliseconds.
- 13:54
Uh, then you're doing some speech-- uh, sorry, text to speech, and then you're going over the network and out the other side. These are real measured numbers, sort of medians.
- 14:06
There's obviously some hand-waving because it's hard to measure everything perfectly, and there's huge amounts of variation. But you're, you're at like six hundred to seven hundred milliseconds optimistically, but realistically in this world.
- 14:17
So that gives you a sort of sense of where we are. Um,
- 14:21
yesterday, we posted a demo on Twitter of really trying to drive all this latency down. Uh, and we have, like, a container that's open source that you can kinda look at and play with and hack on, uh, that gets you to, like, the five hundred sev- to seven hundred milliseconds of, uh, latency.
- 14:40
Hello, I'm Fast Bot. How can I help you today?
- 14:44
Let's play a geography game. Do you wanna give me some geography questions to answer?
- 14:49
That's the spirit. I'll give you a question. What is the world's largest desert?
- 14:54
Oh, I think it's the Sahara.
- 14:56
That's correct. The Sahara Desert is indeed the largest hot desert in the world, covering most of North Africa. Well done. Next question: Which river is the longest in South America?
- 15:07
It's gotta be the, uh, not the Nile, sorry, the Amazon.
- 15:11
No need to apologize. You're absolutely right again. The Amazon River is indeed the longest river in South America and one of the longest rivers in the world. Good job.
- 15:21
Next question: What is the capital city of Australia?
- 15:25
Canberra?
- 15:27
Correct again. Canberra is the capital city of Australia. You're on a roll. Next question: Which mountain range runs along the border between France and Spain?
- 15:36
Pyrenees?
- 15:38
Exactly right. The Pyrenees mountain range does indeed run along the border between France and Spain. You're really showing off your geography skills here. Next question: What is the largest city in Scandinavia?
- 15:50
Visit Stockholm?
- 15:53
I think there might be a slight mistake. Visits.com is not a city- [laughs] ... but rather a travel website.
- 15:58
I was trying to say-
- 15:59
That's correct
- 15:59
... Stockholm.
- 16:01
No worries. I knew what you meant, and you're absolutely right again. Stockholm is indeed the largest city in Scandinavia and the capital of Sweden. You're doing great. Next question: Which city is the largest in Poland?
- 16:14
Kraków?
- 16:16
Close, but not quite. Kraków is a beautiful city in Poland, but it's not the largest. The correct answer is Warsaw, which is the capital and largest city of Poland.
- 16:25
Oh, I feel like I should have known that. If we're gonna talk about capitals, maybe I'll ask you some questions. What's the capital of France?
- 16:33
I'm ready for a challenge. The capital of France is Paris.
- 16:36
Germany?
- 16:38
Not quite. While Germany is a country with-
- 16:40
No, I meant what's the capital of Germany? [laughs]
- 16:43
Good one. The capital of Germany is indeed Berlin.
- 16:47
And what's the capital of Stockholm? I mean, sorry, what's the [laughs] capital of Sweden?
- 16:51
No worries. The capital of Sweden is Stockholm.
- 16:55
What a... So the, the, the idea there was not to do prompt engineering because clearly I could have better, done a better job on prompt engineering, uh, but to kind of really test latency and show all the latency numbers in one place because it's useful to visualize stuff.
- 17:10
Uh, and again, that UI and everything is open source. So the, the big number in the middle, uh, w- is the, the voice-to-voice latency measured on the client side, and then the numbers on the edge are metrics that are being sent from the server that's doing all that work, sort of on the individual pieces, the measurements sent
- 17:27
back, so you can sort of see what the variation is and correlate and kinda get good intuitions about this stuff. Um, the architecture here is, uh, uh, two models, uh, by a company called Deepgram, uh, for the transcription and the voice generation, uh, that are really good compromises between how good they are and how fast they are.
- 17:47
And Deepgram has a hosted service, but they also let you run those models on premises in little Docker containers. Um, and that's Llama 3.8B, I think, because I couldn't quite get 70B to run as fast as I wanted it to, although in theory, uh, that's possible.
- 18:05
Um, and I'll post some links, uh, to this if you want to look at it more.
- 18:11
So because we solved so many problems over and over, uh, we thought it would be great to have an open source framework for this stuff. It- I think we've seen this in other parts of AI landscape, things like LangChain and LlamaIndex are really valuable.
- 18:23
Uh, this is sort of that for real-time and multimodal AI. And this slide probably looks familiar 'cause I stole the list of hard problems from this slide and made a slide that I moved higher up in the talk here for today.
- 18:36
Um, but this is a open source framework called Pipecat. Uh, it's gotten a bunch of traction recently. Uh, it's vendor neutral even though it came out of, uh, work that we've done at Daily early on on this, and we're just really excited about this.
- 18:48
It's super fun to be getting lots of community contributions now. And if you are trying to build really fast multimodal AI stuff, I think it's at least worth taking a look at.
- 18:58
You can build things like conversational bots and speech-to-speech language translation apps and voice controlled agents of various kinds, like, that control your software user interfaces, uh, and real-time vision model stuff like the awesome last presentation is also, like, baked into Pipecat services.
- 19:14
Now, here's all the stuff that's supported in Pipecat today. We're adding stuff all the time. You can add stuff. So if you're interested in building, please hang out with us in the Pipecat Discord.
- 19:26
If you wanna contribute a, a service plugin, please do that. If you wanna be a maintainer for an open source project, that's a lot of fun, ping me. Uh, maintainers are, you know, gold in the open source world.
- 19:37
Uh, we're all always trying to recruit great maintainers. Um, and just last slide about the context here. So this is the Pipecat star rating list, and the day that it went vertical was the GPT-4.0 announcement.
- 19:52
We're gonna get great multimodal models, and they'll be incredibly useful. They'll make building super fast stuff easier and easier. Um, but A, they're not here yet, and B, we're still gonna need orchestration layers for all this stuff.
- 20:05
Um, also the, the, the demo that I showed just a minute ago, uh, that I posted yesterday, is now at 175,000 views on Twitter, so there's more and more and more interest in voice AI, and, uh, we'd love to have people, uh, come build with us. [outro music]