← All AI Engineer talks

AI Engineer Europe 2026

Neil Zeghidour - Voice AI: when is the "Her" moment?

About this talk

Gradium AI CEO Neil Zeghidour examines why contemporary voice agents still fall short of natural, Her-like conversation. Using voice-cloning and Reachy Mini demonstrations, he contrasts cascaded speech-to-text, LLM, and text-to-speech systems with speech-to-speech dialogue, emphasizing half-duplex limitations, simultaneous speech, and Moshi's full-duplex architecture. He also highlights paralinguistic understanding and introduces Gradium Phonon, an on-device text-to-speech model that runs locally on a CPU.

Chapters

  1. 0:00Gradium's voice-model mission and voice-cloning demonstration
  2. 2:50The Her analogy, current voice agents, and Reachy Mini demonstration
  3. 5:42Cascaded voice-agent architectures and half-duplex limitations
  4. 10:46Full-duplex Moshi and paralinguistic understanding
  5. 17:38Gradium Phonon on-device speech synthesis and closing

Talk transcript

  1. 0:00

    [upbeat music] Hi, everyone.

  2. 0:15

    Uh, thanks for being here today. Uh, really happy to, to talk. I think it's the, it's the right time for this talk because we had a lot of great presentations around voice, uh, this morning.

  3. 0:26

    And I want to take a bit, uh, you know, uh, time to reflect where we are in the, uh, modestly where we are right now in the, in the voice, uh, uh, AI, uh, domain, and what is left to do as main challenges.

  4. 0:39

    Uh, quick intro, uh, about Gradient. So we are-- Our mission is to unlock the unrealized potential of voice AI. So basically, we train voice models, uh, speech-to-text, uh, text-to-speech, uh, speech-to-speech, whether it's transformation, translation, uh, speech-to-speech dialogue, and so on and so forth.

  5. 0:56

    We want to be, um, main, uh, model provider, uh, for voice, uh, for everyone building voice agents and voice, uh, solutions. So we are not working on orchestration, we are not working on specific verticals.

  6. 1:08

    We just make, uh, building blocks for people who want to build, uh, voice AI. Um, here, I mean, I can let it be said, uh, uh, by a famous podcaster that I cloned, uh, this morning.

  7. 1:22

    Uh. There should be audio that is output.

  8. 1:32

    Man, have you seen what Gradium is doing with voice cloning? It's kinda crazy. Like, seriously.

  9. 1:37

    Can we put a bit, uh-

  10. 1:38

    You record like ten seconds of your voice. That's it.

  11. 1:39

    Can we have much more volume, please?

  12. 1:44

    Man, have you seen what Gradium is doing with voice cloning? It's kinda crazy. Like, seriously. You record like ten seconds of your voice. That's it. And the system analyzes the tone, the pitch, the accent, all those little quirks that make your voice yours.

  13. 1:58

    Then boom, you type text and it talks back-

  14. 2:01

    Okay, I guess you recognize maybe Joe again. I hope you did. Uh, so basically this is a spinoff from a nonprofit lab we created two years ago, uh, with a funding, uh, uh, from philanthropists, including Eric Schmidt, uh, Rodolphe Se- Rodolphe Seid and Xavier Niel.

  15. 2:15

    The main idea was to create a lab that does, uh, open research. Uh, and, uh, we focused mostly on speech. So we developed Moshi, uh, which was the first, uh, s- uh, speech-to-speech model for conversation, uh, speech-to-speech translation.

  16. 2:28

    Pocket TTS, most recently a CPU model. Uh, and we decided to also create this for-profit structure to, uh, make, uh, products that can be used in production beyond open source.

  17. 2:38

    So the, the goal of the talk is, uh, uh, basically based on the Her movie. So, uh, this has been the most overused, most annoying, uh, analogy, I, I think, in the field.

  18. 2:50

    At the same time, it's extremely relevant, uh, because it's certainly old, and if you look at one of the introduction scenes, so that's when the main character, uh, meets, uh, his AI voice, Samantha, for the first time.

  19. 3:03

    It still sounds like, uh, it was, you know, anticipating, I think, well, what interaction could look like. Oh, what, what do I call you? Do you have a name or...?

  20. 3:14

    Um, yes. Samantha.

  21. 3:17

    Where, where'd you get that name from?

  22. 3:19

    I gave it to myself, actually.

  23. 3:22

    How come?

  24. 3:24

    'Cause I like the sound of it. Samantha.

  25. 3:28

    And, you know, then came the, like, the, the trend of the Her moment. Uh, and we got so many Her moments on, uh, Twitter and in real life. Again, I don't want to be mean to anyone, so I will also make fun of myself.

  26. 3:45

    It's more like, you know, trying to be pragmatic about what was the promise and where we are right now. So this is a very, very recent, uh, demo from, for me, what is the best voice AI company in the world, which is ElevenLabs.

  27. 3:59

    Hello, you're speaking to the government's AI helper. How can I help you today?

  28. 4:04

    I would like to start a new business.

  29. 4:08

    To start a new business, you'll need to choose a business structure and name, complete the registration, and upload required documents.

  30. 4:16

    And this one is a demo that I did, uh, uh, a few months ago with a Richie Mini. So our presenter this morning. Uh, you know, but it, it was very risky, I recognize.

  31. 4:26

    I was very scared when I did it. But, uh, basically putting our, uh, streaming voice models into a, into a Richie Mini.

  32. 4:32

    Action, just say the word.

  33. 4:34

    Okay. Maybe you can do something a bit fun. Um, I want to get, um, you know, to improve my health overall, and so I, I'm looking for a, for a bro to go to the gym.

  34. 4:44

    Can you take the personality and the voice of a gym addict?

  35. 4:50

    Hey, I'm Logan, your gym bro. Let's crush those gains together. You ready to lift, sweat, and feel awesome? I've got your back. No excuses, just results. [laughing]

  36. 4:58

    So, okay, this it's, it's fine. It, it, it sounds more natural than it used to. But in both cases, you know, I mean, we are still not there, right?

  37. 5:07

    Uh, the latency is still quite high. Uh, the ability to handle simultaneous speaking between the user and the system is not there. Uh, intelligence starts to become much better, and I think that's also why we see all this traction around voice agents, because there are agents actually to whom we can give a voice, and they can be

  38. 5:23

    useful. And, um, also this is mostly just a glorified text model with a voice around it. And, uh, so, you know, anything that is not in the text will not be able to be leveraged.

  39. 5:35

    And maybe I could ask one last time to increase a bit the volume. I think that could be even better. Um, and so what does it take to get there?

  40. 5:42

    So basically, we have very nice presentation this morning by Samuel about cascaded systems, so I will go quickly. Speech-to-text, LLM text-to-speech, that's the classic cascade. Uh, in our case, we do, um, uh, streaming speech-to-text, uh, streaming text-to-speech with voice cloning, uh, semantic VAD, the classic stuff.

  41. 6:00

    Um- So latency, you know, getting to a fast, uh, conversation. So we have a very fast TTS, so that's the latency for our TTS compared to a few, uh, other models.

  42. 6:10

    But, you know, that means that just the TTS is still more than two hundred milliseconds, while in a human conversation you need the entire stack of understanding, producing an answer and pronouncing it to be around two hundred milliseconds.

  43. 6:23

    So that will none, none of this will allow to have a conversation that sounds, uh, that sound human. And this is just latency for text conversation. There is no tool call, no actual task that is performed.

  44. 6:37

    Now that's another scene from the- from the movie, uh-

  45. 6:41

    Okay, let's start with your emails. You have several thousand emails regarding LA Weekly, but it looks like you haven't worked there in many years.

  46. 6:47

    Oh, yeah. Okay. So here it just went instantly into through all the emails and gathered information and so on and so forth. Obviously, we are not there at all.

  47. 6:56

    Today, if you have a, a voice agent that is supposed to use a tool, uh, you're going to wait. Uh, the tool call, you know, we are fighting for latency of the TTS, trying to grab ten milliseconds, twenty milliseconds, and then you have a tool call or open router that is going to have a latency between five

  48. 7:11

    hundred milliseconds and four seconds. So in a way, sometimes you feel like we're fighting for something that is not the most relevant anymore. I think now the main bottleneck is becoming the tool call, which is very unpredictable, and we should have models that are resilient to very complex tool calls.

  49. 7:25

    So what we are, uh, you know, one solution to do that is to have fillers. So basically, you know, your LLM, it splits, uh, into two things. So it sends, uh, a tool call and while it waits for getting the result back, it can keep the conversation going in a natural way, and then it retrieves the result

  50. 7:44

    and it tries to insert it back, uh, naturally in the conversation. So just because I always do live demos, I will do a very short, uh, improvised live demo with something that was vibe coded very quickly, so it may be go very wrong.

  51. 7:59

    But the idea here is to have like a just a vibe coded, uh, travel agent and, uh, that can, when I ask it to retrieve, uh, places, it's going to try to find a nice thing to say about the location I'm going to.

  52. 8:11

    Hello. This is Colin from Wanderlust Travel. I see you're looking to book a trip for two people from April tenth to the thirteenth, twenty twenty-six. Where would you like to travel to for your getaway?

  53. 8:23

    Um, I want to go to Tokyo.

  54. 8:26

    Tokyo is such an incredible choice. It's a fascinating mix of ultra-modern skyscrapers and beautiful, peaceful shrines.

  55. 8:35

    I found some fantastic options for you. You could stay at the Fairmont for a serene-

  56. 8:40

    So this needs polishing. The main idea, you know, is that while it was trying to gather things, it doesn't really know how long it will take to get the data.

  57. 8:47

    It tries to say a nice thing about, uh- about the place you're going to. So that's one way to get latency that is going to be more controlled, uh, uh, and more reliable even despite the complexity of the thing.

  58. 9:02

    But then we know that latency w-- of cascaded system is inherently high, right? And so then what I hear a lot from customers, from investors and so on is what about speech-to-speech?

  59. 9:12

    So I think a big confusion is what is speech-to-speech. So speech-to-speech is the idea that now instead of having the three blocks, you only have one that does everything together, right?

  60. 9:20

    So instead of having audios that go in speech-to-text and then LLM text-to-speech, you have a model that takes speech as input and is going to output speech. And that reduces latency a lot, but that's still not a human conversation.

  61. 9:33

    In particular, because every single speech-to-speech model except Moshi is half-duplex. What that means is even the best, uh, speech-to-speech model we could argue, I don't know, maybe that's the advanced voice model for OpenAI or Sesame.

  62. 9:46

    I'm a huge fan of the Sesame, uh, uh, voice, uh, uh, voice model. It's still half-duplex. What it means is that the model is either listening or it's speaking, and it cannot handle the ambiguity of human conversation where you can have overlap between, uh, people speaking on one another.

  63. 10:02

    A lot of things happening. You cough, uh, you know, you just do, "Mm-hmm," which is called backchanneling, and then it breaks completely. Full duplex, that's a human conversation. It varies between cultures and languages.

  64. 10:13

    But for example, in Japanese, it's a sign of politeness and that you are actively listening to do a lot of backchanneling. So you say, "Mm-hmm, uh, uh, uh," constantly when the other person is speaking, and you get up to twenty percent of the time that is overlapped between the people, right?

  65. 10:28

    So that's what makes a conversation human. And so I showed this video in my hotel room this morning just to sh-- you know. That's a speech-to-speech model, but that's not a full duplex model.

  66. 10:37

    And you see how it can, it can become annoying. Hey, how's it going?

  67. 10:41

    Hey, I'm doing really well. Thanks for asking. How about you? How's your day going so far?

  68. 10:46

    Uh, I'm great. Um, you know, I'm preparing a talk about voice AI and how far we are from, um, uh, you know, um, from the Her movie. And, uh, yeah, I'd just like to brainstorm a bit about it with you.

  69. 11:00

    Oh, that's a great topic. Yeah-

  70. 11:02

    Mm-hmm

  71. 11:02

    ... I'd love to help you brainstorm. Are you thinking more-

  72. 11:05

    Yeah.

  73. 11:05

    Exactly. Yeah. I was thinking maybe-

  74. 11:07

    No, no, I didn't mean to interrupt. You know, I was just saying, yeah, like that. You can just keep going. Uh, you know, don't mind me. That's something I typically do.

  75. 11:14

    Ah, no worries at all. Yeah. I was just gonna say we could break it down into a few aspects.

  76. 11:18

    Sure.

  77. 11:19

    Like... Yeah, exactly.

  78. 11:20

    No, no, please stop, stop interrupting. You know, it's called backchanneling. Uh, humans do it all the time. It shows that you're just following the conversation. That shouldn't, you know, like interrupt you in your flow.

  79. 11:29

    Just, just keep going.

  80. 11:31

    Ah, got it. Thanks for letting me know.

  81. 11:33

    No problem.

  82. 11:33

    Yeah. No problem.

  83. 11:34

    Oh, come on. [laughs] Okay. So yeah, basic-- I was a bit mean, right? But, you know, that's my point. It's not, it's not an actual conversation, right? And it can become very annoying in particular.

  84. 11:43

    That's also why, you know, most of the voice AI demos, they are shot in empty, like a quiet room next to the phone and so on. A lot of things can break.

  85. 11:51

    So what we did instead, uh, in our case was, uh, sorry, where am I in my presentation? I'm right here with Moshi. Was, uh, the first full duplex system.

  86. 12:01

    So here you'll see my, uh, co-founder Alex talking to it. Uh, it's almost two-year-old now. I think it's still kind of, uh, aged well because what you'll see is, um, you know, they are going to talk on one another constantly and- It's just fine.

  87. 12:16

    So the planet is Sirius 22. Can you plot a trajectory course to it, please?

  88. 12:21

    Yes, sir.

  89. 12:22

    Okay. How long is it gonna take us to get there?

  90. 12:24

    I've mapped it out.

  91. 12:25

    Okay.

  92. 12:25

    It's approximately five months to get there.

  93. 12:28

    Okay, that's, that's not too bad. Uh, do you think we have all we need on board the ship to start the mission?

  94. 12:33

    Yes, sir. We have everything we need.

  95. 12:35

    Okay.

  96. 12:35

    So, you know, even when the model has guessed what you're going to say, it starts answering before you're done. At the same time, you can talk over it, and it's not ignoring what you're saying.

  97. 12:43

    It's really, you know, s- consider it afterwards and so on. So you, you have what is the most robust, uh, to this day, uh, conversational experience, robust to noise, to a lot of people speaking, and so on and so forth.

  98. 12:55

    But now if we compare it to, um, uh, to the Her movie-

  99. 13:00

    So do you know what I'm thinking right now? Well, I take it from your tone that you're challenging me. Maybe because you're curious how I work? Do you wanna know how I work?

  100. 13:08

    Yeah, actually.

  101. 13:09

    So maybe this was not very clear, but this snippet here, it's the AI understanding that the character is a bit uncomfortable. So that's paralinguistic understanding. You know, it's understanding all the cues that are come from the way people speak.

  102. 13:22

    Technically, that is in Moshi, that is in any speech-to-speech model because this information is not lost. However, if you don't exploit this information, uh, to make your model say relevant things, it's never going to exploit it.

  103. 13:34

    If you train it on the audio version of a instruct dataset, uh, and it's just factual question answering, wh- why would it even try to capture this information? So basically, Moshi, I think we, we saw still the f- only full duplex model.

  104. 13:48

    Uh, recently NVIDIA published a Personal Plex model based on it. Um, what was great is still the, the flow of it is just honestly impossible to match, I guess.

  105. 13:57

    Um, it's conversational and, uh, and very robust. At the same time, the model was very stupid. So basically it was, you know, just useless. You could talk for, for a few minutes, and then it was a bit pointless, you know.

  106. 14:07

    Because it was not an agent, it had no tool call, no ability to do anything. Uh,

  107. 14:13

    it's impossible to use in production something that is, has no observability. You don't know if the... It's very hard to detect if someone said something that should be not accepted and so on and so forth.

  108. 14:23

    And there was no real, uh, paralinguistic, uh, understanding. So

  109. 14:29

    the main takeaway for me is, uh, we know that this nature of interactive model that are going to be full duplex, that are going to be, um, really the, that will be the way to get to

  110. 14:43

    an interaction that is as natural as you would have with a human. But as long as we don't, are ab- we are not able to give to this kind of very natural sounding, uh, models the same level of reliability, intelligence, and personalization as cascaded systems, I don't see a path towards, uh, uh, them, you know, replacing cascaded

  111. 15:04

    systems. So I used to be really at war against cascaded systems. I think they are so practical and so convenient, that the main challenge, honestly, I think this, we solved that with, uh, Moshi.

  112. 15:14

    Uh, anybody who implements it and train it on better data will have something that is, sounds just indistinguishable from humans. But all of this is going, all the challenge is going to happen here.

  113. 15:23

    Uh, the last point is the scalability. So now let's say you have the best, uh, speech-to-speech model, okay? Uh, it solves ev- all the stuff that I've talked to, talked about.

  114. 15:34

    If we take again the analogy from the movie, um, you will talk to it maybe several hours a day, or it will always be on because, you know, it's on your computer when you work and you asynchronously ask things to it and so on and so forth.

  115. 15:45

    Um, I'm not going to mention the cost of the API of our competitors, but voice is very expensive. Everybody in this room probably knows it. Uh, the voice mode of most hyperscalers is run at a loss.

  116. 15:57

    It's a gigantic multimodal model, and they lose money every time they, they, you use it. But it's, you know, it's kind of a marketing thing, and so on. It's fine.

  117. 16:05

    But now if we want to make it an actual profitable product that people are using at massive scale, just not going to work. And, uh, in particular, anyone who tried developing a consumer app with voice, uh, realize that LLM now is almost nothing in terms of cost.

  118. 16:20

    All the bill is TTS. Speech-to-text is very cheap as well. Uh, diarization is affordable. Uh, TTS is really what is going to consume most of the... And we, I saw people burning their fundraising in TTS bills, and they don't even get the opportunity to get their user base to grow.

  119. 16:36

    So another aspect is privacy. The more you're going to open to your AIs, the more, um, you know, you'll want it to be more controlled and private and feel more comfortable that things are not shared publicly.

  120. 16:48

    In particular, uh, we see now people with Mythos being afraid that any, any single database was going to be hacked in a few months. And so, you know, you will be more comfortable if all your private data is, uh, is local.

  121. 17:00

    And so to solve that, our first step is, uh, Gradium Phonon. Uh, it's a on-device TTS. So on-device means a lot of things. For some people, on-device means it runs on a gamer GPU.

  122. 17:11

    For us, on-device means it runs on the smartphone CPU. And so it's a very small model, uh, less than one million parameter. And for its size, it works quite well.

  123. 17:20

    It's better than all the existing, uh, uh, on-device models. CoCo is, is a good one, but it doesn't have voice cloning. And I'm out of time, so I would just play, uh, a short, uh, demo if I can.

  124. 17:34

    Yes.

  125. 17:38

    Uh, geez, Morty, stop looking for a signal. Gradium Phonon runs locally on the CPU, which means high fidelity neural speech right here without those intergalactic cloud government hacks.

  126. 17:47

    It is about simplicity with no servers and no waiting, just a smooth, quiet performance that stays on the device with local processing and total privacy.

  127. 17:55

    What else?

  128. 17:57

    Hmm, local processing. That sounds like the machine is making its own donut. Wait, if the CPU is doing all the work, can it sprinkle some-

  129. 18:04

    Yeah, so this runs on a smartphone CPU, which means that you can use that to power, you know, any kind of voice application without paying a single, uh, cent of, uh, API, uh, fee.

  130. 18:14

    So we opened the private beta for this, uh, model. Uh, the goal for us is to allow people to create consumer apps with voice, basically. And that they are able to scale the usage, uh, without having to lose money on the, on API.

  131. 18:28

    So the conclusion, uh, the path forward is, uh, for us, I have a str- strict opposition to some of our competitors say that voice is a commodity now. I think it's completely false.

  132. 18:39

    Uh, voice is very challenging. The last mile is going to be the most difficult, uh, to solve. And, uh, for us, it's really about science, uh, and engineering. Uh, and it will get Her to us, to the Her movie, sorry.

  133. 18:52

    So you can use us on gradium.ai, and if you want to join us, uh, to bring Her to life, yes. No, so I, I [laughs] now I'm using this analogy.

  134. 19:01

    No, that's if you want to, to, to work on very exciting, uh, voice models, you can apply and, and, uh, and join us. Thanks a lot for your attention. [clapping] [upbeat music]