AI Engineer World's Fair 2026
Voice agents with Realtime Video — Sidney Primas, LemonSlice
About this talk
LemonSlice cofounder and CTO Sidney Primas presents real-time visual agents intended to approach an avatar Turing test. He demonstrates an interactive historical avatar and single-image characters, explains how LemonSlice adds an API-based visual layer to independently supplied language and voice models, and discusses model orchestration, emotional expression, evaluation, and future operating costs.
Chapters
- 0:00LemonSlice and the avatar Turing test
- 1:34Historical-avatar demonstration and visual AI interaction
- 5:58Single-image avatar demo and API architecture
- 14:57Model harness and emotionally expressive real-time video
- 22:17Avatar evaluation, cost questions, and closing
Talk transcript
- 0:00
[upbeat music] My name is Sidney. Uh, I am the CTO, uh, and founder of Lemon Slice.
- 0:19
And Lemon Slice is on a mission, uh, to break the avatar Turing test. What we mean by this is, uh, making an avatar, uh, that is indistinguishable from a human on a video call.
- 0:39
Uh, and all of this is, of course, making the avatar photorealistic. Uh, but there's actually a long tail of technical problems, uh, that we care a lot about that I'll be talking about today, uh, that we're planning to solve.
- 0:53
Uh, and those problems are things like getting the emotions right, uh, getting, uh, the object interactions right, getting the, like, microexpressions right, uh, and even figuring out, uh, an internal state for these avatars so that they feel, uh, like a real human being.
- 1:12
So what I'm gonna do in this... Oh.
- 1:15
What I'm gonna do in this talk is talk about our technical bet to solving the avatar Turing test, uh, what we've built today, uh, and the technical, the hard technical problems we've solved today, uh, and our bet on, like, how we think this problem will be solved in the long, long term.
- 1:34
Uh, but before we dive into that, I wanna give you guys a sense, uh, for what our product is and how it works. Uh, so actually yesterday was a, was a pretty big and exciting day for us.
- 1:45
Uh, we had a big, big launch yesterday. Uh, we had a partnership with Microsoft, uh, that brought an ex-US president to life, uh, Teddy Roosevelt. Uh, and they built his presidential library, and inside of the presidential library, uh, they did a replica of the Oval Office, and you can walk in and you can have a full interaction,
- 2:07
uh, with Teddy Roosevelt in real time. Uh, so definitely not, uh, Turing test solved yet, but, like, a big milestone, uh, towards trying to solve, solve that problem. Uh, and the coolest part of all of this was, uh, we had a special guest.
- 2:24
Uh, but hate, hate him or love him, uh, he showed up. Uh, so we had President Trump, uh, who came to the opening, uh, of this ceremony, uh, and interacted, uh, with our avatar.
- 2:37
So, uh, we actually got some footage, real footage, uh, not, not fake, uh, AI-generated footage of Trump, uh, interacting. So let me, let me show you guys this real quick.
- 2:48
Do you consider the Panama Canal your greatest achievement?
- 2:54
Well, well. [laughs] Right then, yes. Well, then. The canal stands as one of my proudest battles, no question. But greatness is a strange thing. It isn't always the biggest or boldest job that matters most.
- 3:10
Panama Canal showed what America could achieve if we held steady and acted fast when the world dragged its feet. That said, I measure my greatest work by the lives improved.
- 3:22
Parks set aside, food and drugs made safe, the square deal given to all, not just to a few. Still, when I stood in the mud watching those steam shovels, knowing ships would pass through, changing the world's map forever, I felt I'd left a mark that would last.
- 3:39
Perfect. Uh, so you guys get a sense, uh, of, of the experience. Um, the funny thing is he was scheduled to be there for, you know, a quick minute, one interaction, and he actually stayed for ten minutes.
- 3:54
He left, he left the room, uh, and then came back again to talk to Teddy, Teddy Roosevelt even more. Um, so, um, you could see there, like, the avatar is there.
- 4:04
It's a full body avatar. It has hands, it does actions, it does movement. Um, and, uh, all of this is obviously, uh, done in real time. Uh, I'll talk a little bit more about, like, the why
- 4:16
we think this is an important problem, uh, later in the talk.
- 4:21
Um, but a big part of it is actually related to this interaction. Um,
- 4:27
humans, like, struggle to, to, to focus and pay attention. Uh, it's, it's often hard for a lot of people to spend a bunch of time reading. Um, and even voice, it's hard to really lock in and understand what's happening.
- 4:39
Uh, so we think, like, we're biologically wired to best, like, understand things, uh, when there's a visual component as well. So our bet here is, in the long term, we think most interactions between AI a-and humans will have a visual, visual layer, uh, and we're building that visual layer.
- 5:02
Um, so now I wanna talk about our approach to building these avatars. It's a very different approach than what most other avatar companies use.
- 5:17
Uh, essentially what we do is we take these world models, and we focus them on humans. Um, and the reason we take this bet is even though it's harder to get the initial model working, uh, it's harder to train the model, it's harder to deploy the model, once you have a model, you get all of these nice
- 5:39
emergent properties that we all know about, uh, where you just very easily can solve things like full body movement, like object interactions, uh, like movements i-in-in the scene, uh, all the way down to the mi-micro, microexpressions as well and emotions.
- 5:58
And all of those kind of don't come for free, but come more for free than when you use the other approaches that other people use. So let me guys, let me show you guys real quick the product in action so you guys can get a sense.
- 6:14
Um, so let me expand this a little bit.
- 6:32
How is it going? I am here to listen and help with anything you need. We can chat about something in particular or play a game if you would like.
- 6:41
So the lip sync might be off here a little bit, uh, but that's just because of the AV setup here. Uh, but as you can see, um, we have actions.
- 6:52
Uh, they can wave. Uh, they really can do any actions. Uh, we have movement. We have hands. Um, and the way, the way this works is it, it just can use any single image to create the avatar.
- 7:05
Uh, and so, yes, you can have photorealistic. Uh, you also can have Pixar. You also can have cartoons. Uh, it just literally takes, uh, a single image. And, um,
- 7:17
show you this. Oh, no. Perfect. Uh, and I'll, I'll show you kind of a quick interaction here as well, and then, and then we can talk about how we built this.
- 7:33
Um, here, within the same video call and on the same kind of, like, inference setup, uh, we can easily change clothe. We can change the scene she's in, and
- 7:46
you can talk to it too. I'm not gonna do it here.
- 7:47
You sound like you might be trying to say story. Would you like to tell me a story or hear one?
- 7:52
Yeah. So you can see here too, like, that there's physics, uh, on the earrings. They move. Uh, the water moves. Uh, so, like, the, the video model has a good understanding of, of physics and, like, can use that to, like, uh, make the avatar feel more realistic.
- 8:10
Perfect. Now, the way this is used, just for, for kind of everybody's information, is we are mostly the API layer. Uh, so we provide an API. Uh, people bring their own LLM.
- 8:23
Uh, people bring their own usually, like, voices. Um, and then we're the, the visual layer o-on top of it. Uh, so people use this for, um, language learning. People use it for, like, AI sales call.
- 8:38
Basically anything today where you have a voice agent, you now also kind of have a vis-- video agent if it's on your screen or on your computer.
- 8:48
Great. So now let's talk about how we built this today. Um, this is our approach. Um, and can't talk about all the technical details, uh, but, uh, I'm gonna try to go give you enough information to make it interesting.
- 9:05
Um, so what we do is we get started with training our own, uh, video DiT model. Um, the big thing here for us that really matters, uh, is the audio.
- 9:15
Uh, the audio turns out to be very important for getting emotions right and the facial expressions right. Uh, people care about that a lot. Uh, so we really focus on, um, getting great audio data that we can train on, uh, and then actually also, um, getting the audio encoders very right.
- 9:37
Um, like, most audio encoders today are trained on basically, uh, audiobooks, which is very monotone, very simple, uh, don't have a lot of emotions. So if you wanna have a very expressive model, uh, you can't use those audio co-- encoders, and you actually have to spend a lot of time, like, getting the audio embeddings right so that
- 9:57
the video model is super expressive. Uh, so, so that's kind of like step one for us. Uh, and this is basically just a video model. Um, you know, people can call this a world model, but, but it's basically a, a, a, a video model, uh, that understands the physics of the world.
- 10:11
And now the big challenge is: How do you take that video model, and how do you make it real-time and interactive? Um, and so the way that works is the, the interactivity part is interesting.
- 10:22
So, uh, usually video models are bidirectional, so they can look into the past, but they actually also can look into the future and, like, what's about to come. And they basically generate videos, uh, all at the same time by looking at all of the, the latents that, that are being generated.
- 10:39
Uh, for us, you can't do that. You only can look into the past. Uh, and so here on this, like, Make Interactive, you can see that, uh, we basically train a model with an attention mask so that the model can only look into the past.
- 10:54
So when you do inference, uh, it never can see the future because the future doesn't exist because, like, you haven't given it those inputs yet. So for example, it doesn't see audio in the future.
- 11:03
It only see what has, has been said in the past. And then, uh, once you have that, we make it real-time. Um, and there's a bunch of things that goes in here, but the biggest thing to speed up, uh, video generation is, uh, usually you spend, like, a bunch of steps denoising these video models.
- 11:24
Uh, so, like, let's say thirty steps. You spend thirty steps, like, removing the noise, uh, to generate the beautiful, beautiful videos. Um, and what we need to do is go from, like, thirty steps, bring it down to one step.
- 11:36
So you basically just in a single step go from, like, pure noise to a pure video model to make it, make it real-time. Uh, so those are the, the, the two big things that, like, make it interactive and real-time.
- 11:46
Uh, now the hard part here, like, what makes this difficult is this problem here on the left, uh, which is called error accumulation, uh, which is a problem anybody in, like, real-time video generation and world model is familiar with.
- 12:00
Uh, I'll show you this video, and then we talk about it briefly.
- 12:03
And I'm a professor-
- 12:04
Oh
- 12:04
... in ophthalmology at Stanford School of Medicine.
- 12:08
Uh, so you can see, I mean, that's a very bad example, uh, o-o-o-of error accumulation. Uh, but you can see as the video continues to generate, uh, more error is introduced.
- 12:18
Uh, and the reason for that is because since you only can look backwards, you're looking actually backwards at, at, at, uh, videos that you've previously generated. But each video you generate or each video block you generate has some error in it.
- 12:33
So now you're looking in the past, you're looking at the error, you're adding more error to it, and then just the error compounds over time. Um, and this is especially hard because, I mean, ideally, these, these video models are endless.
- 12:45
Like, the, the Teddy ava-avatar is generating continuously non-stop frame by frame for eight hours straight with, like, no reset throughout the entire process. Uh, we have another one that's gonna be generating for sixteen hours straight.
- 12:58
So it's a very hard problem to not have any error accumulation over long periods of time. And, um,
- 13:06
we came up with a new way to solve this problem, uh, that is different to the best of our knowledge than what everybody else, uh, does today. And so I can't talk about that yet, uh, but because of that, we can basically generate these very long videos, uh, I mean, you guys can, can try it out, uh,
- 13:23
that essentially have no error accumulation, no noticeable error accumulation in them. Um,
- 13:30
great. Um, and so the last two hard problems, uh, we worked on to enable this is, uh, hardware optimization. Um,
- 13:44
there's, there's a bunch of stuff that goes into this. The big thing here is, uh, making it low cost enough, uh, so you actually can use this. Like, usually when you generate an AI video, it's five seconds, and you just share that.
- 13:58
Uh, we have to generate minutes and hours of AI, AI video and still have it be cost-effective. Uh, and so the cool, cool thing here is we've been able to make the models small enough and efficient enough so that the costs are about the same as a voice model, which is crazy to me.
- 14:16
'Cause, like, think about the, a voice model, how much data is streamed there compared to a video model. Like, it- it's much more pixel heavy in a video model, uh, than with voice.
- 14:27
Uh, and the costs, costs are about similar. Um, so the cool thing there is this enables us to do, uh, consumer use cases, and a lot of our customers are consumer companies that, uh, basically use the video for consumer and, like, entertainment applications.
- 14:44
And then the final thing to mention here is the model harness. Um, I feel like the model harness is something that is often overlooked but is actually super important and super hard.
- 14:57
Uh, like, getting the model harness right, uh, is a huge technical challenge for us, and I feel like a lot of our value actually in, like, productizing this is in the model harness.
- 15:06
And the way to think about it is you just have a bunch of separate threads. Uh, all, all of it is managing real data streaming through our system, and you have basically a bunch of stuff you do on a GPU and a bunch of stuff you do on a CPU, and you have to orchestrate this perfectly in
- 15:23
a way that, like, the video always remains real-time. There's never any stutter that happens inside of the video. Um, and this is especially hard when you have things like interrupts, you have queues, you're buffering data, you have to clean the queues.
- 15:35
And so, like, getting this orchestration right at production, at scale has been a ton of work. And honestly, I feel like over time, uh, y- a lot more of the value, uh, uh, of, like, the things we build will be in, like, figuring out the model harness.
- 15:49
I, I think it's especially true for, like, any real-time applications. Um, cool. So this is, this is our technical approach of, like, commercializing these, uh, world models for avatars.
- 16:01
Uh, let me talk a little bit about, like, what we're actually working on actively now. Um, so
- 16:08
the big thing we wanna enable next is having more of an emotional engine, uh, that's driving these avatars, uh, having them be more, more aware of what's actually happening inside of the conversation and then be able to react to what's happening.
- 16:23
So react means, uh, the right e- the right emotions at the right time for the right duration, and same, same for the actions. Um, and, um, today when you're, when you interact with these avatars, you can try, it still feels a little bit-- I mean, it definitely still feels, um, uh, awkward.
- 16:42
Uh, and I think a big part of that is, uh, they don't emotionally react to you. They, they're not listening to you. They're not, like, emoting in the right way when they're talking.
- 16:51
Um, and so this is what we're trying to solve. And so let me show you guys some of, um, kind of the, the new videos we're able to create, uh, with, with basically, basically better emotional control.
- 17:05
Um, so this is a real-time video.
- 17:07
Oh my God. Goal. Goal. We won the game. I can't believe it. We won. Oh my God. Oh my God. Oh my God.
- 17:20
Oh, shoot. So you can see, like, a lot, like, the emotions just matter a lot. Uh, they make a huge difference in, like, connecting with this avatar. You feel, you feel much more connected with that avatar.
- 17:31
Um, so that's one thing. And then, uh, the next-- this is the next generation model. Uh, not in real-time yet, uh, but there's just a lot more-
- 17:39
I think the most important thing is just to take a moment and breathe. Life gets so busy, you know?
- 17:45
Do you think you'll stay there for a while?
- 17:46
Yeah. I really think I need to stay right here for a long time.
- 17:50
So there, there'll be a lot more-
- 17:51
I think it's just to take-
- 17:52
... basically natural interaction.
- 17:53
Decided what to do about the apartment?
- 17:56
I don't know.
- 17:56
Sorry. Uh, there'll be a n- a lot more natural interactions with- ... uh, themselves and then also with their environment. And the in- the, the big thing here is, like, our model can already do this today.
- 18:07
Um, it, it has the capabilities to do this. It's just not controllable enough to, like, make it real time with the conversation, uh, and not deterministic enough to, to make it useful with the conversation.
- 18:17
So a big part of our work here is, like, making these actions actually controllable at the right, right time. And to do that, we're building this emotion engine that is just literally predicting based on, like, the audio input from the avatar and the text input for what the avatar is gonna say, what the action they should be
- 18:35
doing at this moment in time. Um, so target launch for this next model is,
- 18:44
um, in, in basically one to two month is what we're targeting.
- 18:48
Um, awesome.
- 18:50
So where do you want to be tonight?
- 18:51
Oh, shoot. All right. Great. Um, so that's where we are today. Um, I just wanna talk briefly about... Let's see what the time is. Almost done. Um, I wanna talk briefly about where we see this heading.
- 19:08
So I strongly believe that in the end, uh, there'll be a single model, um, that is the EQ layer for AI. Uh, the way you can think about this model is the model will literally take in directly, uh, the user video and the user audio, so directly the u- th- you know, the user talking to you like
- 19:32
you are in a FaceTime, uh, and put out the avatar video and audio. And do all of that in an end-to-end model, in, in a single model end to end.
- 19:42
Um, internally inside of the model, it will do the audio understanding, generate what it's gonna say, uh, model its own internal state, like own internal emotional state, uh, and then based on all of that, produce an output
- 20:01
video and audio. And all of this will be trained in one model. Like, this will happen. Uh, we're already seeing, like, early papers that are, like, doing proof, proofs of concepts around this.
- 20:11
What we're not saying is that this EQ model will be very intelligent. Uh, it'll be very-- it will have very high EQ, and it'll be very good at, like, interacting with people.
- 20:22
Uh, but there will be a separate model, uh, that will drive the IQ, that will basically give this EQ model input to, like, do all the magical things that AI can do today, like do the tool calling, do the deep thinking, do all the intelligent stuff.
- 20:38
Uh, but there will be an end-to-end layer up front. And so that's, that's our bet. Like, we're working towards, towards building that. Uh, and, um, we can talk more about, like, some of the benefits of this approach, but I feel strongly that within two or three years, you'll be seeing these kinds of end-to-end EQ models coming on
- 20:57
the market. Um, that's it. Um, well, maybe take some questions if there's any questions. Um,
- 21:07
yeah, go ahead.
- 21:08
Um, so obviously, like EQ could probably define, like, thousands of different configurations of emotion for human beings.
- 21:16
Mm-hmm.
- 21:17
So obviously, like, the inputs are pretty clear, like video and audio, and, like, you can process that some way by the model. How do you think about, like, the outputs of the EQ model and, I guess, the granularity at which you will be able to control that?
- 21:31
In the future? Well, I think in the future, in the long term, uh, it's gonna be more of an internal state, uh, where the model will, will have an internal state in latent space that, like, defines what the emotional state is or the goals are of the model, and they will be more implicit, so they won't be
- 21:50
human interpretable necessarily. Um, today, today, basically the way we think about actions and emotions is basically words. If you can describe it in words, you can generate those actions and emotions.
- 22:03
So I guess, like, the state of, like, the EQ that you expect the model to hold would be, I guess, what is the ideal state of the avatar in response to whatever you're decoding from the input model?
- 22:17
Exactly. [crosstalking] Exactly. And track that over time.
- 22:20
Yeah.
- 22:20
Right? Like, have an internal state that tracks that over, over some history of time. But yes, exactly.
- 22:27
You guys are trying to pass the, the Turing test, the avatar Turing test.
- 22:31
Yes.
- 22:32
Uh, I just wondered if you passed it with your first example. Did Trump know that wasn't Roosevelt? [laughs]
- 22:38
Uh, there's a big debate about that. [laughs] Uh, Twi- Twitter is debating. Gavin Newsom said, uh, that he definitely doesn't know. Uh, and, you know, uh, he got tricked by, like, the ghost of, of, uh, of Teddy Roosevelt. [laughs]
- 22:55
Similar question, like, so the avatars will be like humans. What is the way that companies can make the distinction between the real human followers on the platform?
- 23:06
You're asking about what the way method is or the test is? Yeah. No, that's a really good point. We haven't done this yet, but we're, uh, in the process of figuring out our own version of the Turing test for these avatars, which will just include real people.
- 23:23
Uh, and, uh, we're planning to run that this year. Um, and we won't pass it this year, but it'll just be cool for us to track it over time to see when we can actually solve the Turing test, uh, and ideally allow other people to also
- 23:40
go through the test as well. But yeah, we plan to r- we plan to create a test and then publish around, around the test.
- 23:50
Oh, yeah.
- 23:52
Are you also planning to, uh, create a digital twin extending to, uh, this to make-
- 24:08
Uh, not for us. Like, that's not what we plan to do. Um, maybe people can use our technology to do that, uh, but that's not our, our goal. Uh, but you could imagine, uh, people using our technology, giving it the right context, having the right hardness around it, uh, and then having the nice, like, IQ brain around
- 24:30
it to do exactly what you're saying. But, like, we, we don't wanna solve that problem.
- 24:42
How do you imagine the cost per hour of using this scaling over the next few years?
- 24:49
Is it gonna be a big challenge for using it in the real world or is it very expensive currently?
- 24:54
Yeah. Um, it's still-- We would still like it to be cheaper. Um, the good news is it, it is-- Okay, almost done. Last question. Um, the good news is it is surprisingly-- Like, going into this, we didn't know how cheap this would be.
- 25:07
Uh, I've been very surprised at how inexpensive it is. Again, like, the cost of this is at the same level as an audio model, uh, in terms of what we charge for it.
- 25:16
Um, and we're all hoping for this costs to go down. Part, part of it is for consumer, you just need very low costs. Uh, and then part of it is the, as the cost goes down, we can do higher resolution, which will help us a lot.
- 25:29
Um, but I guess the way we're thinking about it is there's, like, algorithm improvements, uh, and hardware improvements. Um, and the-- What has been true for some time now is things have been improving on all vectors faster than expected, and we're just hoping that continues, [chuckles] uh, basically.
- 25:49
I think there'll also be very cool, uh, architectural updates to, to move to more of, like, a token approach instead of a diffusion approach that will make video, like, this type of video generation way cheaper.
- 26:00
Um, that's it for me. Uh, thank you guys, uh, for showing up. Um, we, if you guys are-- We are hiring. If you guys are interested in, like, solving these types of technical problems, come talk to me afterwards.
- 26:13
Uh, or if you're just interested in this at all, come talk to me afterwards. Appreciate it, guys. [audience applauding] [upbeat music]