← All AI Engineer talks

AI Engineer Europe 2026

Neil Zeghidour - Voice AI: when is the "Her" moment?

Neil Zeghidour· Co-founder & CEO, Gradium19:27

Read the talk

Voice AI’s remaining gap: from convincing speech to useful conversation

A cloned voice can sound convincing while the conversation still fails. Neil Zeghidour traces the remaining problems through latency, overlapping speech, agent capabilities and on-device synthesis.

From a talk by Neil Zeghidour

From a voice recording to a synthetic speaker

What does it take to reproduce a voice—and how much further is that from building an assistant people can talk to naturally? Gradium supplies the model building blocks: speech recognition, speech synthesis, speech transformation, translation and spoken dialogue. Neil Zeghidour distinguishes that work from orchestration and applications for particular industries. The intended customers are developers assembling voice agents and other voice products.

“Who is Gradium?” slide with a team photograph, mission statement, STT, TTS and S2S text, and a colorful dotted graphic.
Gradium introduces its mission and voice-model focus.

The opening demonstration makes the synthesis problem tangible. A synthetic voice, which Zeghidour identifies as a Joe Rogan clone, describes recording roughly ten seconds of speech, capturing tone, pitch, accent and vocal quirks, then speaking newly supplied text. That reference length is a claim in the demonstration, not a measured cloning benchmark. The striking result is the separation between vocal identity and content: the recognizable voice can deliver words supplied afterward.

Gradium grew out of a nonprofit research lab backed by philanthropists including Eric Schmidt, Rodolphe Saadé and Xavier Niel. Its open research included Moshi, speech-to-speech translation and the CPU-oriented Pocket TTS. The commercial structure was created to turn that research lineage into products usable in production. Convincing synthesis is therefore the starting point for the talk, not its definition of a finished voice assistant.

0:260:39
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:26 · section reference included

A natural voice is only part of the interaction

In the introductory scene from Her, the human asks his new assistant what to call her. Samantha supplies a name, explains that she chose it herself, and says she likes its sound. The exchange feels spontaneous because the voice, timing and responses work together. Despite the overuse of the film as an industry analogy, it remains a useful reference for the interaction being promised.

Two contemporary demonstrations show both progress and the remaining distance. An ElevenLabs government helper answers a request to start a business with steps covering structure, naming, registration and documents. Zeghidour then shows his own experiment putting streaming voice models into a Reachy Mini robot. Asked to become a gym enthusiast, the robot adopts the name Logan and an energetic gym-bro voice, offering encouragement to lift and train. One demonstration supplies useful procedural information; the other changes voice and personality on request.

Embedded video showing two seated people with a small white robot between them on an ai-PULSE stage.
An onstage robot demonstration appears under “In real life …”.

Both sound more natural than earlier voice systems, but response delays and simultaneous speech remain problems. Better language-model intelligence now gives developers useful agents to put behind a voice. Yet a system whose reasoning depends on a text transcript cannot use vocal information that the transcript leaves out. Natural-sounding output does not by itself produce natural conversation.

2:382:50
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:38 · section reference included

The latency budget extends beyond speech synthesis

The conventional cascade runs speech-to-text → LLM → text-to-speech. Streaming recognition and synthesis let components begin processing before an entire utterance or response is complete; voice cloning controls the output voice, while semantic voice activity detection helps determine when the user has finished. These are useful building blocks, but fast components do not automatically make the combined interaction fast enough.

Zeghidour reports that TTS alone still takes more than 200 milliseconds in the comparison he presents, against a conversational target of around 200 milliseconds for the complete response. The human comparison needs a precise interpretation: turn-taking research describes short gaps between speakers, with response planning overlapping incoming speech, rather than all comprehension and production happening serially inside that gap. Even the simpler machine latency comparison here excludes tool calls and actual task execution.

Another Her scene raises the bar: Samantha inspects thousands of emails and immediately notices that correspondence concerns a job the user left years ago. Real agents must wait for external work. Zeghidour cites tool-call or OpenRouter delays of 500 milliseconds to four seconds, compared with efforts to save 10–20 milliseconds in TTS. These are illustrative ranges without a specified workload or measurement protocol. The engineering problem becomes resilience to unpredictable tool completion, not merely shaving time from speech generation.

One approach separates the tool operation from the speech that keeps the interaction moving:

  1. Send the tool request.
  2. Continue speaking naturally while the result is pending.
  3. Incorporate the returned information when it becomes available.

The challenge is that the model does not know in advance how much conversational space it needs to fill. The speech must lead naturally into the result rather than sound like an unrelated prerecorded waiting message.

The live travel-agent demonstration follows this pattern. Colin from Wanderlust Travel asks where two people would like to go for their April 10–13, 2026 trip. Zeghidour chooses Tokyo. The assistant first describes the city’s skyscrapers and peaceful shrines, then announces accommodation options and begins mentioning the Fairmont. The destination commentary occupies the interval before the options arrive. Zeghidour calls the prototype unfinished: this manages perceived waiting and conversational continuity; it does not make the underlying retrieval finish sooner.

5:426:00
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:42 · section reference included

Speech-to-speech does not guarantee simultaneous conversation

Replacing the cascade with a single speech-to-speech model removes boundaries between separate recognition, reasoning and synthesis components. Speech goes in and speech comes out, reducing latency. But that architectural choice does not settle whether the system can listen while it speaks.

Interaction modeListening and speakingConsequence
Half duplexAlternates between themUser speech can force a turn change
Full duplexSupports both simultaneouslyOverlap can remain part of the conversation

Zeghidour characterizes prominent alternatives, including OpenAI’s advanced voice offering and Sesame’s voice model, as half duplex at the time of the talk. That is his assessment of contemporary systems, not a permanent classification of their products. The distinction matters because a cough or brief acknowledgment need not mean the user wants the assistant to stop.

Backchanneling is speech that signals attention without taking the floor. Its use varies across languages and cultures. Zeghidour gives Japanese conversation as an example, where frequent acknowledgments can express politeness and active listening, and cites overlap reaching 20% of conversational time in that context. A voice system that treats every acknowledgment as an interruption makes an ordinary listening behavior disruptive.

Slide with green and orange timelines labeled USER and AI for half-duplex, and ALICE and BOB for full-duplex.
Half-duplex turn-taking contrasts with overlapping full-duplex conversation.

The hotel-room demonstration exposes that failure directly. Zeghidour asks an assistant to brainstorm a talk about voice AI and Her. As it starts answering, he supplies brief acknowledgments: mm-hmm, yeah and sure. Its responses become interrupted or restarted instead of developing the requested ideas. He explains that he is only signaling attention and asks it to keep going, but the pattern continues. Instructions about conversation cannot compensate for an interaction mechanism that repeatedly yields the floor at the wrong time.

The example is deliberately demanding, as Zeghidour acknowledges, but the behavior it tests is commonplace. It also helps explain why a polished demonstration recorded beside a phone in a quiet room is insufficient evidence of conversational robustness. Overlap, noise and less controlled surroundings expose additional ways an interaction can break.

9:029:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:02 · section reference included

Moshi keeps listening while answering

Moshi supplies the contrasting demonstration. In a recording Zeghidour describes as nearly two years old, his co-founder Alex role-plays a spaceship conversation. He asks for a course to Sirius 22, asks how long the journey will take, and checks whether the ship has everything needed to begin. Moshi responds while the human continues speaking and acknowledging its answers. The fictional navigation exchange demonstrates conversational timing, not an executed navigation tool.

Two behaviors distinguish this from simply ignoring interruptions. The model can anticipate where an utterance is going and begin answering before the user finishes. At the same time, speech arriving during its answer remains available for subsequent responses. Continuing to speak and continuing to listen are compatible operations. Zeghidour also describes Moshi as robust to noise and multiple speakers, although this clip is an illustration of overlap rather than a comparative robustness benchmark.

11:5112:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:51 · section reference included

Retaining vocal cues is not the same as using them

The next Her exchange introduces a different requirement. Samantha interprets the human’s tone as a challenge and connects it to curiosity about how she works. Zeghidour reads the scene as recognition of discomfort: the assistant responds to how something is said, not only to the literal words. This is paralinguistic understanding.

Speech-to-speech models retain acoustic information that a text transcript can discard. Availability, however, does not guarantee useful behavior. If training consists of spoken versions of factual instruction-following examples, the correct answer may be identical regardless of tone. Such data gives the model little reason to exploit vocal cues. Learning to respond to those cues requires examples in which the manner of speaking changes what a relevant response should be.

Zeghidour then points to NVIDIA’s PersonaPlex, which builds on Moshi. Its full-duplex design also makes clear that Moshi is not literally the only full-duplex model; the useful distinction is the interaction architecture and its descendants. Moshi’s fluid conversation was a strong foundation, but Zeghidour describes the original model as lacking enough intelligence to remain useful after a short exchange. It had neither tool calls nor the ability to perform tasks.

Production introduces further requirements:

  • Observability: Developers need to inspect behavior and detect speech that should not be accepted.
  • Paralinguistic behavior: Preserved audio cues must actually influence responses; the original Moshi did not provide meaningful understanding of this kind.
  • Reliability, intelligence and personalization: Natural interaction must meet the practical expectations already served by cascaded systems.

These requirements explain why conversational flow alone is insufficient to replace an existing agent stack.

Zeghidour consequently revises his earlier opposition to cascades: their practicality and convenience matter. He believes the full-duplex approach, trained on better data, can deliver humanlike conversational flow. But that remains his projection; the harder replacement test is whether such models can also provide the capabilities and controls that make an agent dependable.

12:5513:00
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:55 · section reference included

Always-on voice changes the economics

Suppose the conversational and agent problems are solved. An assistant resembling Samantha might then be used for hours each day, or remain available throughout work while the user occasionally asks it to do something. That makes scalability a product requirement rather than an optimization to postpone until after voice quality is solved.

Zeghidour argues that large multimodal voice services can be expensive enough to operate at a loss, though he supplies no financial evidence for that claim. His more concrete concern comes from consumer-app developers: in the situations he describes, text LLMs, speech recognition and diarization are relatively inexpensive, while TTS dominates the bill. He recounts seeing teams spend their fundraising on synthesis before they have a chance to grow their user base. The relevant constraint is sustained usage cost, not whether a short demo is affordable.

Persistent use also invites more personal disclosure. As users entrust an assistant with more of their lives, control over where information resides becomes more consequential. Zeghidour invokes fears surrounding Mythos and future database attacks as a motivation for keeping private data local; those fears are not an established forecast of universal compromise. The practical motivation is to reduce the amount of sensitive information that must leave the device.

15:2315:34
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:23 · section reference included

Moving synthesis onto a smartphone CPU

Gradium Phonon is presented as a first step toward local deployment. Here, on-device means running TTS on a smartphone CPU, not requiring a gaming GPU. Zeghidour describes it as having fewer than one million parameters; Gradium’s April evaluation instead lists approximately 100 million for Phonon 20260318. Those specifications conflict, so the smaller figure should not be treated as settled.

Zeghidour claims strong quality relative to existing on-device models and contrasts Phonon’s voice cloning with Kokoro, which he describes as lacking that capability. The universal superiority claim is broader than the vendor’s published comparison supports: the April evaluation covers particular synthesis tests and alternatives, not every on-device model or smartphone performance condition. The important product combination is local CPU execution with personalized speech.

“Gradium Phonon Real-Time Inference on CPU” slide with CPU inference, personalization and performance boxes beside a table highlighting Phonon.
Gradium Phonon’s CPU inference features and benchmark comparison.

The accompanying audio demonstration cycles through synthetic character voices discussing local processing, including a voice addressing Morty and another joking about donuts. The clips illustrate expressive synthesis, while their script promotes privacy and independence from servers. Local TTS can keep synthesis on the device; it does not establish privacy for an entire application if other components still send data elsewhere.

Running synthesis on the smartphone avoids a per-request TTS API charge, which is the scaling advantage Zeghidour emphasizes as he announces the private beta. It is not cost-free deployment: the launch announcement describes fixed licensing costs by device and model type. The intended change is to let consumer-app usage grow without a corresponding stream of synthesis API charges.

That leaves a substantive last mile. Voice is not a solved commodity merely because generated speech sounds good. Useful everyday interaction still requires science and engineering across conversation, agent behavior and deployment. Zeghidour closes by returning to the Her analogy and directing builders to Gradium: reaching that experience means making natural interaction dependable and affordable enough to use beyond the demonstration.

17:0017:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

17:00 · section reference included

Resources

From the talk

  • Research explaining fast conversational transitions and how response planning overlaps listening.

  • Technical account of contextual speech synthesis, prosody evaluation and conversational modeling limitations.

  • Evaluating PhononArticle

    April benchmark results and methodology for Phonon's synthesis accuracy and voice similarity.

Updates since the talk

Read the complete timestamped transcript
  1. 0:00

    [upbeat music] Hi, everyone.

  2. 0:15

    Uh, thanks for being here today. Uh, really happy to, to talk. I think it's the, it's the right time for this talk because we had a lot of great presentations around voice, uh, this morning.

  3. 0:26

    And I want to take a bit, uh, you know, uh, time to reflect where we are in the, uh, modestly where we are right now in the, in the voice, uh, uh, AI, uh, domain, and what is left to do as main challenges.

  4. 0:39

    Uh, quick intro, uh, about Gradient. So we are-- Our mission is to unlock the unrealized potential of voice AI. So basically, we train voice models, uh, speech-to-text, uh, text-to-speech, uh, speech-to-speech, whether it's transformation, translation, uh, speech-to-speech dialogue, and so on and so forth.

  5. 0:56

    We want to be, um, main, uh, model provider, uh, for voice, uh, for everyone building voice agents and voice, uh, solutions. So we are not working on orchestration, we are not working on specific verticals.

  6. 1:08

    We just make, uh, building blocks for people who want to build, uh, voice AI. Um, here, I mean, I can let it be said, uh, uh, by a famous podcaster that I cloned, uh, this morning.

  7. 1:22

    Uh. There should be audio that is output.

  8. 1:32

    Man, have you seen what Gradium is doing with voice cloning? It's kinda crazy. Like, seriously.

  9. 1:37

    Can we put a bit, uh-

  10. 1:38

    You record like ten seconds of your voice. That's it.

  11. 1:39

    Can we have much more volume, please?

  12. 1:44

    Man, have you seen what Gradium is doing with voice cloning? It's kinda crazy. Like, seriously. You record like ten seconds of your voice. That's it. And the system analyzes the tone, the pitch, the accent, all those little quirks that make your voice yours.

  13. 1:58

    Then boom, you type text and it talks back-

  14. 2:01

    Okay, I guess you recognize maybe Joe again. I hope you did. Uh, so basically this is a spinoff from a nonprofit lab we created two years ago, uh, with a funding, uh, uh, from philanthropists, including Eric Schmidt, uh, Rodolphe Se- Rodolphe Seid and Xavier Niel.

  15. 2:15

    The main idea was to create a lab that does, uh, open research. Uh, and, uh, we focused mostly on speech. So we developed Moshi, uh, which was the first, uh, s- uh, speech-to-speech model for conversation, uh, speech-to-speech translation.

  16. 2:28

    Pocket TTS, most recently a CPU model. Uh, and we decided to also create this for-profit structure to, uh, make, uh, products that can be used in production beyond open source.

  17. 2:38

    So the, the goal of the talk is, uh, uh, basically based on the Her movie. So, uh, this has been the most overused, most annoying, uh, analogy, I, I think, in the field.

  18. 2:50

    At the same time, it's extremely relevant, uh, because it's certainly old, and if you look at one of the introduction scenes, so that's when the main character, uh, meets, uh, his AI voice, Samantha, for the first time.

  19. 3:03

    It still sounds like, uh, it was, you know, anticipating, I think, well, what interaction could look like. Oh, what, what do I call you? Do you have a name or...?

  20. 3:14

    Um, yes. Samantha.

  21. 3:17

    Where, where'd you get that name from?

  22. 3:19

    I gave it to myself, actually.

  23. 3:22

    How come?

  24. 3:24

    'Cause I like the sound of it. Samantha.

  25. 3:28

    And, you know, then came the, like, the, the trend of the Her moment. Uh, and we got so many Her moments on, uh, Twitter and in real life. Again, I don't want to be mean to anyone, so I will also make fun of myself.

  26. 3:45

    It's more like, you know, trying to be pragmatic about what was the promise and where we are right now. So this is a very, very recent, uh, demo from, for me, what is the best voice AI company in the world, which is ElevenLabs.

  27. 3:59

    Hello, you're speaking to the government's AI helper. How can I help you today?

  28. 4:04

    I would like to start a new business.

  29. 4:08

    To start a new business, you'll need to choose a business structure and name, complete the registration, and upload required documents.

  30. 4:16

    And this one is a demo that I did, uh, uh, a few months ago with a Richie Mini. So our presenter this morning. Uh, you know, but it, it was very risky, I recognize.

  31. 4:26

    I was very scared when I did it. But, uh, basically putting our, uh, streaming voice models into a, into a Richie Mini.

  32. 4:32

    Action, just say the word.

  33. 4:34

    Okay. Maybe you can do something a bit fun. Um, I want to get, um, you know, to improve my health overall, and so I, I'm looking for a, for a bro to go to the gym.

  34. 4:44

    Can you take the personality and the voice of a gym addict?

  35. 4:50

    Hey, I'm Logan, your gym bro. Let's crush those gains together. You ready to lift, sweat, and feel awesome? I've got your back. No excuses, just results. [laughing]

  36. 4:58

    So, okay, this it's, it's fine. It, it, it sounds more natural than it used to. But in both cases, you know, I mean, we are still not there, right?

  37. 5:07

    Uh, the latency is still quite high. Uh, the ability to handle simultaneous speaking between the user and the system is not there. Uh, intelligence starts to become much better, and I think that's also why we see all this traction around voice agents, because there are agents actually to whom we can give a voice, and they can be

  38. 5:23

    useful. And, um, also this is mostly just a glorified text model with a voice around it. And, uh, so, you know, anything that is not in the text will not be able to be leveraged.

  39. 5:35

    And maybe I could ask one last time to increase a bit the volume. I think that could be even better. Um, and so what does it take to get there?

  40. 5:42

    So basically, we have very nice presentation this morning by Samuel about cascaded systems, so I will go quickly. Speech-to-text, LLM text-to-speech, that's the classic cascade. Uh, in our case, we do, um, uh, streaming speech-to-text, uh, streaming text-to-speech with voice cloning, uh, semantic VAD, the classic stuff.

  41. 6:00

    Um- So latency, you know, getting to a fast, uh, conversation. So we have a very fast TTS, so that's the latency for our TTS compared to a few, uh, other models.

  42. 6:10

    But, you know, that means that just the TTS is still more than two hundred milliseconds, while in a human conversation you need the entire stack of understanding, producing an answer and pronouncing it to be around two hundred milliseconds.

  43. 6:23

    So that will none, none of this will allow to have a conversation that sounds, uh, that sound human. And this is just latency for text conversation. There is no tool call, no actual task that is performed.

  44. 6:37

    Now that's another scene from the- from the movie, uh-

  45. 6:41

    Okay, let's start with your emails. You have several thousand emails regarding LA Weekly, but it looks like you haven't worked there in many years.

  46. 6:47

    Oh, yeah. Okay. So here it just went instantly into through all the emails and gathered information and so on and so forth. Obviously, we are not there at all.

  47. 6:56

    Today, if you have a, a voice agent that is supposed to use a tool, uh, you're going to wait. Uh, the tool call, you know, we are fighting for latency of the TTS, trying to grab ten milliseconds, twenty milliseconds, and then you have a tool call or open router that is going to have a latency between five

  48. 7:11

    hundred milliseconds and four seconds. So in a way, sometimes you feel like we're fighting for something that is not the most relevant anymore. I think now the main bottleneck is becoming the tool call, which is very unpredictable, and we should have models that are resilient to very complex tool calls.

  49. 7:25

    So what we are, uh, you know, one solution to do that is to have fillers. So basically, you know, your LLM, it splits, uh, into two things. So it sends, uh, a tool call and while it waits for getting the result back, it can keep the conversation going in a natural way, and then it retrieves the result

  50. 7:44

    and it tries to insert it back, uh, naturally in the conversation. So just because I always do live demos, I will do a very short, uh, improvised live demo with something that was vibe coded very quickly, so it may be go very wrong.

  51. 7:59

    But the idea here is to have like a just a vibe coded, uh, travel agent and, uh, that can, when I ask it to retrieve, uh, places, it's going to try to find a nice thing to say about the location I'm going to.

  52. 8:11

    Hello. This is Colin from Wanderlust Travel. I see you're looking to book a trip for two people from April tenth to the thirteenth, twenty twenty-six. Where would you like to travel to for your getaway?

  53. 8:23

    Um, I want to go to Tokyo.

  54. 8:26

    Tokyo is such an incredible choice. It's a fascinating mix of ultra-modern skyscrapers and beautiful, peaceful shrines.

  55. 8:35

    I found some fantastic options for you. You could stay at the Fairmont for a serene-

  56. 8:40

    So this needs polishing. The main idea, you know, is that while it was trying to gather things, it doesn't really know how long it will take to get the data.

  57. 8:47

    It tries to say a nice thing about, uh- about the place you're going to. So that's one way to get latency that is going to be more controlled, uh, uh, and more reliable even despite the complexity of the thing.

  58. 9:02

    But then we know that latency w-- of cascaded system is inherently high, right? And so then what I hear a lot from customers, from investors and so on is what about speech-to-speech?

  59. 9:12

    So I think a big confusion is what is speech-to-speech. So speech-to-speech is the idea that now instead of having the three blocks, you only have one that does everything together, right?

  60. 9:20

    So instead of having audios that go in speech-to-text and then LLM text-to-speech, you have a model that takes speech as input and is going to output speech. And that reduces latency a lot, but that's still not a human conversation.

  61. 9:33

    In particular, because every single speech-to-speech model except Moshi is half-duplex. What that means is even the best, uh, speech-to-speech model we could argue, I don't know, maybe that's the advanced voice model for OpenAI or Sesame.

  62. 9:46

    I'm a huge fan of the Sesame, uh, uh, voice, uh, uh, voice model. It's still half-duplex. What it means is that the model is either listening or it's speaking, and it cannot handle the ambiguity of human conversation where you can have overlap between, uh, people speaking on one another.

  63. 10:02

    A lot of things happening. You cough, uh, you know, you just do, "Mm-hmm," which is called backchanneling, and then it breaks completely. Full duplex, that's a human conversation. It varies between cultures and languages.

  64. 10:13

    But for example, in Japanese, it's a sign of politeness and that you are actively listening to do a lot of backchanneling. So you say, "Mm-hmm, uh, uh, uh," constantly when the other person is speaking, and you get up to twenty percent of the time that is overlapped between the people, right?

  65. 10:28

    So that's what makes a conversation human. And so I showed this video in my hotel room this morning just to sh-- you know. That's a speech-to-speech model, but that's not a full duplex model.

  66. 10:37

    And you see how it can, it can become annoying. Hey, how's it going?

  67. 10:41

    Hey, I'm doing really well. Thanks for asking. How about you? How's your day going so far?

  68. 10:46

    Uh, I'm great. Um, you know, I'm preparing a talk about voice AI and how far we are from, um, uh, you know, um, from the Her movie. And, uh, yeah, I'd just like to brainstorm a bit about it with you.

  69. 11:00

    Oh, that's a great topic. Yeah-

  70. 11:02

    Mm-hmm

  71. 11:02

    ... I'd love to help you brainstorm. Are you thinking more-

  72. 11:05

    Yeah.

  73. 11:05

    Exactly. Yeah. I was thinking maybe-

  74. 11:07

    No, no, I didn't mean to interrupt. You know, I was just saying, yeah, like that. You can just keep going. Uh, you know, don't mind me. That's something I typically do.

  75. 11:14

    Ah, no worries at all. Yeah. I was just gonna say we could break it down into a few aspects.

  76. 11:18

    Sure.

  77. 11:19

    Like... Yeah, exactly.

  78. 11:20

    No, no, please stop, stop interrupting. You know, it's called backchanneling. Uh, humans do it all the time. It shows that you're just following the conversation. That shouldn't, you know, like interrupt you in your flow.

  79. 11:29

    Just, just keep going.

  80. 11:31

    Ah, got it. Thanks for letting me know.

  81. 11:33

    No problem.

  82. 11:33

    Yeah. No problem.

  83. 11:34

    Oh, come on. [laughs] Okay. So yeah, basic-- I was a bit mean, right? But, you know, that's my point. It's not, it's not an actual conversation, right? And it can become very annoying in particular.

  84. 11:43

    That's also why, you know, most of the voice AI demos, they are shot in empty, like a quiet room next to the phone and so on. A lot of things can break.

  85. 11:51

    So what we did instead, uh, in our case was, uh, sorry, where am I in my presentation? I'm right here with Moshi. Was, uh, the first full duplex system.

  86. 12:01

    So here you'll see my, uh, co-founder Alex talking to it. Uh, it's almost two-year-old now. I think it's still kind of, uh, aged well because what you'll see is, um, you know, they are going to talk on one another constantly and- It's just fine.

  87. 12:16

    So the planet is Sirius 22. Can you plot a trajectory course to it, please?

  88. 12:21

    Yes, sir.

  89. 12:22

    Okay. How long is it gonna take us to get there?

  90. 12:24

    I've mapped it out.

  91. 12:25

    Okay.

  92. 12:25

    It's approximately five months to get there.

  93. 12:28

    Okay, that's, that's not too bad. Uh, do you think we have all we need on board the ship to start the mission?

  94. 12:33

    Yes, sir. We have everything we need.

  95. 12:35

    Okay.

  96. 12:35

    So, you know, even when the model has guessed what you're going to say, it starts answering before you're done. At the same time, you can talk over it, and it's not ignoring what you're saying.

  97. 12:43

    It's really, you know, s- consider it afterwards and so on. So you, you have what is the most robust, uh, to this day, uh, conversational experience, robust to noise, to a lot of people speaking, and so on and so forth.

  98. 12:55

    But now if we compare it to, um, uh, to the Her movie-

  99. 13:00

    So do you know what I'm thinking right now? Well, I take it from your tone that you're challenging me. Maybe because you're curious how I work? Do you wanna know how I work?

  100. 13:08

    Yeah, actually.

  101. 13:09

    So maybe this was not very clear, but this snippet here, it's the AI understanding that the character is a bit uncomfortable. So that's paralinguistic understanding. You know, it's understanding all the cues that are come from the way people speak.

  102. 13:22

    Technically, that is in Moshi, that is in any speech-to-speech model because this information is not lost. However, if you don't exploit this information, uh, to make your model say relevant things, it's never going to exploit it.

  103. 13:34

    If you train it on the audio version of a instruct dataset, uh, and it's just factual question answering, wh- why would it even try to capture this information? So basically, Moshi, I think we, we saw still the f- only full duplex model.

  104. 13:48

    Uh, recently NVIDIA published a Personal Plex model based on it. Um, what was great is still the, the flow of it is just honestly impossible to match, I guess.

  105. 13:57

    Um, it's conversational and, uh, and very robust. At the same time, the model was very stupid. So basically it was, you know, just useless. You could talk for, for a few minutes, and then it was a bit pointless, you know.

  106. 14:07

    Because it was not an agent, it had no tool call, no ability to do anything. Uh,

  107. 14:13

    it's impossible to use in production something that is, has no observability. You don't know if the... It's very hard to detect if someone said something that should be not accepted and so on and so forth.

  108. 14:23

    And there was no real, uh, paralinguistic, uh, understanding. So

  109. 14:29

    the main takeaway for me is, uh, we know that this nature of interactive model that are going to be full duplex, that are going to be, um, really the, that will be the way to get to

  110. 14:43

    an interaction that is as natural as you would have with a human. But as long as we don't, are ab- we are not able to give to this kind of very natural sounding, uh, models the same level of reliability, intelligence, and personalization as cascaded systems, I don't see a path towards, uh, uh, them, you know, replacing cascaded

  111. 15:04

    systems. So I used to be really at war against cascaded systems. I think they are so practical and so convenient, that the main challenge, honestly, I think this, we solved that with, uh, Moshi.

  112. 15:14

    Uh, anybody who implements it and train it on better data will have something that is, sounds just indistinguishable from humans. But all of this is going, all the challenge is going to happen here.

  113. 15:23

    Uh, the last point is the scalability. So now let's say you have the best, uh, speech-to-speech model, okay? Uh, it solves ev- all the stuff that I've talked to, talked about.

  114. 15:34

    If we take again the analogy from the movie, um, you will talk to it maybe several hours a day, or it will always be on because, you know, it's on your computer when you work and you asynchronously ask things to it and so on and so forth.

  115. 15:45

    Um, I'm not going to mention the cost of the API of our competitors, but voice is very expensive. Everybody in this room probably knows it. Uh, the voice mode of most hyperscalers is run at a loss.

  116. 15:57

    It's a gigantic multimodal model, and they lose money every time they, they, you use it. But it's, you know, it's kind of a marketing thing, and so on. It's fine.

  117. 16:05

    But now if we want to make it an actual profitable product that people are using at massive scale, just not going to work. And, uh, in particular, anyone who tried developing a consumer app with voice, uh, realize that LLM now is almost nothing in terms of cost.

  118. 16:20

    All the bill is TTS. Speech-to-text is very cheap as well. Uh, diarization is affordable. Uh, TTS is really what is going to consume most of the... And we, I saw people burning their fundraising in TTS bills, and they don't even get the opportunity to get their user base to grow.

  119. 16:36

    So another aspect is privacy. The more you're going to open to your AIs, the more, um, you know, you'll want it to be more controlled and private and feel more comfortable that things are not shared publicly.

  120. 16:48

    In particular, uh, we see now people with Mythos being afraid that any, any single database was going to be hacked in a few months. And so, you know, you will be more comfortable if all your private data is, uh, is local.

  121. 17:00

    And so to solve that, our first step is, uh, Gradium Phonon. Uh, it's a on-device TTS. So on-device means a lot of things. For some people, on-device means it runs on a gamer GPU.

  122. 17:11

    For us, on-device means it runs on the smartphone CPU. And so it's a very small model, uh, less than one million parameter. And for its size, it works quite well.

  123. 17:20

    It's better than all the existing, uh, uh, on-device models. CoCo is, is a good one, but it doesn't have voice cloning. And I'm out of time, so I would just play, uh, a short, uh, demo if I can.

  124. 17:34

    Yes.

  125. 17:38

    Uh, geez, Morty, stop looking for a signal. Gradium Phonon runs locally on the CPU, which means high fidelity neural speech right here without those intergalactic cloud government hacks.

  126. 17:47

    It is about simplicity with no servers and no waiting, just a smooth, quiet performance that stays on the device with local processing and total privacy.

  127. 17:55

    What else?

  128. 17:57

    Hmm, local processing. That sounds like the machine is making its own donut. Wait, if the CPU is doing all the work, can it sprinkle some-

  129. 18:04

    Yeah, so this runs on a smartphone CPU, which means that you can use that to power, you know, any kind of voice application without paying a single, uh, cent of, uh, API, uh, fee.

  130. 18:14

    So we opened the private beta for this, uh, model. Uh, the goal for us is to allow people to create consumer apps with voice, basically. And that they are able to scale the usage, uh, without having to lose money on the, on API.

  131. 18:28

    So the conclusion, uh, the path forward is, uh, for us, I have a str- strict opposition to some of our competitors say that voice is a commodity now. I think it's completely false.

  132. 18:39

    Uh, voice is very challenging. The last mile is going to be the most difficult, uh, to solve. And, uh, for us, it's really about science, uh, and engineering. Uh, and it will get Her to us, to the Her movie, sorry.

  133. 18:52

    So you can use us on gradium.ai, and if you want to join us, uh, to bring Her to life, yes. No, so I, I [laughs] now I'm using this analogy.

  134. 19:01

    No, that's if you want to, to, to work on very exciting, uh, voice models, you can apply and, and, uh, and join us. Thanks a lot for your attention. [clapping] [upbeat music]