← All AI Engineer talks

AI Engineer World's Fair 2025

Building Conversational AI Agents - Thor Schaeff, ElevenLabs

About this talk

ElevenLabs developer-experience engineer Thor Schaeff leads a hands-on workshop on multilingual conversational AI agents, demonstrating regional accents, language-specific text-to-speech voices, and the Voice Library. Attendee questions explore agent model selection, latency and RAG, multilingual transcription limits, voice-creator moderation, and pronunciation of specialized acronyms; a colleague identified as Paul assists the workshop.

Chapters

  1. 0:00Workshop introduction and ElevenLabs developer experience
  2. 11:54Accent comprehension and multilingual demonstrations
  3. 17:45Text-to-speech voices, Voice Library, and regional accents
  4. 24:35Hindi agent demonstration
  5. 35:21Audience questions: pricing, agent orchestration, latency, and RAG
  6. 44:47Multilingual limitations, language learning, and live moderation
  7. 54:04Custom vocabulary, acronym pronunciation, and closing

Talk transcript

  1. 0:00

    [electronic music] All right.

  2. 0:16

    Hello, everyone. I hope you have enough space. [laughing] [chuckles]

  3. 0:22

    Um, no, thanks so much for joining. Um, we're building multilingual conversational AI agents. I know it's a bit of a mouthful, um, but yeah, hopefully we'll, we'll get something going.

  4. 0:36

    Um, yeah, so different than on the poster, we're not from Evelyn Labs, we're from ElevenLabs. Um, have... Can I get a quick show of hands, have you heard of ElevenLabs?

  5. 0:49

    Okay, everyone has heard of it. That's great. Um, so we can, we can do some nice things today. Maybe look at some, um, new stuff maybe you've, you haven't, uh, played with yet.

  6. 1:02

    So this is me. Uh, I'm Thor, uh, here on the, the right side of the screen. I work on developer experience at ElevenLabs, so, you know, a lot of it is kind of the conversational AI, um, you know, agents platform.

  7. 1:18

    And I also have my colleague, uh, Paul here with me. So if you have kinda any questions, you know, throughout the workshop, um, later on, we'll be floating around and, you know, happy to kind of, uh, um, you know, answer your questions.

  8. 1:32

    So feel free to just put your hand up and we'll be, uh, floating around later. Um, yeah, Paul works with me on developer experience. So if you have any feedback as well, uh, you know, on documentation, examples, developer experience, or generally just the product, um, do let us know.

  9. 1:52

    Um, yeah, also this, you can, uh, scan the QR code to get the slides. So there's a couple of resources, um, that are linked in the slides, as well as there's a form, um, where you can, uh, fill in your email address if you want to get, uh, some credits.

  10. 2:07

    So we can give you, um, you know, a couple credits for the next three months to sort of, you know, play around with ElevenLabs and kind of try this out.

  11. 2:15

    Um, so yeah, feel free to just scan this and then, you know, kind of save it, uh, for later on. There's the resources linked in there as well.

  12. 2:26

    Cool. Um, yeah, also if you're building with ElevenLabs, I recommend you follow our ElevenLabs Devs Twitter account. Um, this is specifically for, you know, updates in terms of API versions, client libraries.

  13. 2:39

    So anything, you know, if you're a developer building with ElevenLabs, that's a good place to follow and kind of, um, you know, be in the loop with what is happening.

  14. 2:49

    And then, um, yeah, I mentioned, so if you, if you scanned the QR code earlier, um, there's also the link, so you can tap on the QR code as well to open the form.

  15. 3:00

    Uh, and if you just fill in your, um, email address, we will send you after this workshop. Um, so in the workshop you can get started with kind of the free account.

  16. 3:09

    Uh, that should be, should be plenty. But then, um, you know, we'll, we'll give you a coupon code. Uh, we'll send it to you via email for kind of the next three months to, to try it out.

  17. 3:20

    Cool. And as I mentioned, yeah, the resources in the slide. So we'll, um, you know, kind of tap into these later on, uh, and then can kind of, uh, you know, get started building.

  18. 3:32

    We can see as well kind of what folks are building. Um, you know, since we're looking at kind of multilingual, um, conversational AI agents, do you just wanna shout out kind of the, the languages that, you know, you're specifically looking to unlock with conversational AI?

  19. 3:49

    Anyone, you know, anything else than English? Any other languages? Portuguese. Portuguese. Very nice. Uh, are you looking for like a Brazilian Portuguese accent or... Cool. Yeah, we can, we can look at that.

  20. 4:04

    So we've got Portuguese. Um, any other languages?

  21. 4:08

    Spanish. Spanish, yes. We got some Spanish in there as well. That's good. So Portuguese, Spanish.

  22. 4:15

    I've eaten a product from Hungary. Hungarian. Hungarian. Very good. Um, actually, do we have Hungarian right now? We don't yet, but I think hopefully soon. So we're working on the version three of the multilingual models, and I do think we'll, we'll need to double-check, uh, not that I promised something wrong, but we might be able, maybe not

  23. 4:37

    today, but maybe in a couple of weeks we can give you Hungarian. That's great. Uh, okay.

  24. 4:44

    Any other... Mandarin? Hindi. Hindi, yeah. Um, a huge population, uh, that speaks Hindi. Then again, India has fifty-plus languages, I believe. So we're, we're working on some...

  25. 4:59

    adding some additional, uh... Currently, we have Hindi and Tamil. Um, so we're working on some additional languages there. Okay. So we have a good mix that we can play around with.

  26. 5:08

    Um, cool. Yeah, if there's any other languages later, we can, you know, kind of explore those as well. Uh, maybe one language that probably doesn't get spoken about, um, often enough, uh, is barking.

  27. 5:23

    Um, actually, if you, uh, want to build applications for the nine hundred million dogs that are out there, uh, I actually looked this up, uh, you can use our model.

  28. 5:33

    So this is our most recent launch, um, uh, two months ago now, I believe, uh, the, the newest model we launched. So maybe we can, can have a little listen.

  29. 5:42

    Uh, the dachshund is my favorite actually. [dog barking] Um, the golden retriever. [dog barking]

  30. 5:51

    More. And Chihuahua. [dog barking] Yeah, that one's a bit, bit feisty. But um, yeah, if you're building application for dogs, that might be a great use case for you. Um, did, did anyone see this launch like two months ago?

  31. 6:10

    No? Yeah, you saw it? Um, yeah, the f-unfortunate thing is we launched it, uh, 1st of April actually, and so everyone thought it was an April Fool's. Um, so the timing was a bit unfortunate on that one.

  32. 6:23

    Um, but as you can see, you know, Text to Bark, it's very real. Uh, no. In fact, it was an April Fool's, so be careful when you play this to your dog.

  33. 6:33

    They might get offended, um, because the context, we can't guarantee that it translates. But the actual sounds that you're hearing are generated by our sound effects model. So we do have a, a model.

  34. 6:47

    If you go to the app, um, there is this, uh, sound effects model here, um, and you can actually... [keyboard clicking] This sound effect was also created. [laughs]

  35. 6:58

    But, so truck reversing, uh, maybe that's a good one. You can click Generate. Um, and so basically, what we do is we generate kind of four different samples for you that you can use.

  36. 7:10

    So, um, you know, not, not so interesting for your conversations per se, but like, you know, if you're creating video games, for example. I don't know if anyone, um, is doing that.

  37. 7:21

    We might be having some, uh, internet. I hope we have internet, no?

  38. 7:28

    Okay, the internet is looking okay. So yeah, not sure what's happening here. But so for example, this was like a [cowbell ringing] drum cowbell I, I generated recently. Um, actually, we do have, um, if you Google ElevenLabs sound, um, board, uh, we built this recently, which was-- which is pretty cool.

  39. 7:51

    Um, so you can kind of loop, uh, it's, it's basically like a drum machine, uh, and you can get kind of the [rain falling]

  40. 8:02

    sound effects. Uh, the drums are pretty nice as well, and, and, you know, they are mapped to like the, the keys on the keyboard. [drum music] So you can, you can play like...

  41. 8:16

    So that's pretty cool, and then you can add kind of new sound effects and stuff. You know, just to give you an idea of some of the things, um, that we do.

  42. 8:26

    Now, obviously here, you know, we're talk- we're talking about conversational AI agents. Um, and so specifically, if we kind of look at the different components that are involved in building conversational AI agents, we have the user, um, who is speaking, you know, some language.

  43. 8:43

    Uh, and so we need to transcribe that speech into text. We then feed that into a large language model, which is kind of acting as the brain, um, of, you know, our agent.

  44. 8:54

    So in our case, we currently don't build any intelligence models, so we partner with kind of the existing large language models, um, providers, so like your GPT-4O, uh, your Google Gemini, what have you.

  45. 9:08

    Um, and then this large language model, which will act as kind of the brain of your agent, will generate a text output, uh, and then we basically stream that text output back into speech.

  46. 9:21

    So that is roughly kind of the, the pipeline that we have for building conversational AI agents. Now, with that, there is a bunch of system tools that are built in.

  47. 9:30

    So for example, we have this language detection system tool, which facilitates the language, um, switching, um, that, that we're gonna look at in a bit. Uh, we also have, you know, function calling, tool calling.

  48. 9:41

    So, um, you know, if you need to give access to kind of specific functionality to your agent, you can, you can do that here as well. And then, um, you know, there's different approaches.

  49. 9:52

    So like if you've seen OpenAI real-time, for example, so this doesn't actually go through text, right? So it goes sound token to sound token, which has some benefits, but what we've seen, uh, you know, for deploying conversational AI agents at scale, uh, and really understanding, you know, what is kind of happening.

  50. 10:10

    So if you're going sound token to sound token, you're kind of flying blind a little bit, so you know, you're trusting the, the model that sort of it actually repr-- replies intelligently.

  51. 10:21

    Um, whereas if you're going through text, you can, uh, have, you know, kind of better monitoring and sort of understanding what's, what's going on, um, in your conversation. So while we're also exploring, you know, kind of sound-to-sound for kind of the conversational, uh, AI agents, for now, what we found works best is sort of this, this pipeline

  52. 10:41

    that we've built. Um, and we deploy kind of all these, uh, models very close to each other to kind of bring down the latency as, as much as possible.

  53. 10:51

    Cool. And so now, now we can look at kind of the, the individual, um, sort of components within that. So for example, for speech-to-text, uh, so this is, uh, actually the, the most recent model we actually launched, um, you know, that wasn't an April Fool's.

  54. 11:07

    And so this is our, um, speech-to-text model, so our, uh, automatic speech recognition model, so ASR. Um, and this works, you know, kind of benchmark leading across ninety-nine different languages, um, at the moment.

  55. 11:23

    So what this does, and you know, you can see sort of the functionality that's built in there. There's, you know, word-level timestamps. There's speaker diarization. Um, there's audio event tagging.

  56. 11:34

    So if you want in your transcript, you know, coughing, laughing, sort of some of these, you know, audio tags in there, you can enable that as well. Uh, and it's all kind of within, you know, structured API, uh, responses, which is really nice.

  57. 11:49

    So what you can see here is, for example, you know, we have a conference call.

  58. 11:54

    Quick check-in. [REDACTED:location] is a mess. Time to fix it.

  59. 12:00

    Totally. Some of those potholes could swallow a small car.

  60. 12:03

    Or a very brave skateboarder.

  61. 12:06

    We start next week.

  62. 12:08

    So what you can see here is, as we're kind of playing this audio, we see that the model recognizes the different speakers, uh, and kind of tags them as speaker one, speaker two.

  63. 12:18

    Um, and then we have the word-level timestamp. So you see as I play this- Jonas, four-week timeline?

  64. 12:24

    Yep, unless the concrete throws a t-

  65. 12:27

    So we can, we can highlight the different words kind of, you know, word level, um, here. So that's really, really useful. Um, and, you know, obviously this is available through, uh, the API.

  66. 12:39

    So actually one thing, um, I did, and you can, um, try this out yourself. This is just, uh, available kind of for free to sort of demo it. Um, so if you're using Telegram, I've built kind of this little Telegram bot where you can forward voice messages, um, or videos, uh, to, you know, the Telegram bot and

  67. 13:00

    then you-- it automatically identifies, okay, what language is that? And it gives you back the, the transcript, um, of that message. So, you know, you know this all too well.

  68. 13:10

    You're sitting in a meeting, and your grandmother sends you a voice message, and you don't know, oh, is this urgent? Is this important? So what you can do is you can forward it to the bot, uh, and very quickly you will get the transcript back.

  69. 13:22

    So, um, if you wanna try this out, you know, you can, you can do that. So you can see it here. I was actually--

  70. 13:30

    I hope you, you recognize this, uh, if you were paying attention. Um, anyone recognize this?

  71. 13:40

    You were just saying that.

  72. 13:41

    Yes, I was just saying that. Uh, so I, I was just recording a voice message on my phone. Thank you. Uh, someone was paying attention. That's great. Um, and so, you know, I was just recording a voice message, and then, uh, I get the transcript back.

  73. 13:55

    Now, the cool thing as well, um, so I, I, I actually live in, in Singapore. So if you spend time in Singapore, you might have heard kind of the Singlish, right?

  74. 14:04

    Which is sort of the Singapore English.

  75. 14:07

    Go, go to hell, ah, this person. Bastard, you know. You-- I asked for a plastic bag. You must put the thing inside the plastic bag for me, right? And she never.

  76. 14:15

    She just put the plastic bag, throw the plastic bag on, on the table. Yeah, so throw-

  77. 14:18

    Anyone understand what's going on? Do we got any Singaporeans in the house? No?

  78. 14:26

    It's not that easy. Even after six years in Singapore, um, I still sometimes struggle with that. So what we can do is we can forward it to our transcription bot.

  79. 14:36

    Um, we can see, okay, it was received. It's transcribing it now. And, uh, yeah, "Go to hell. Bastard, you know. I asked for a plastic bag. You must put the thing inside the plastic bag for me."

  80. 14:48

    So you can, you can see these are kind of the problems we have in, uh, Singapore. Now, another accent that might be a bit challenging, um, Scottish.

  81. 14:59

    All right.

  82. 15:00

    What are some things I need to know about Scotland?

  83. 15:02

    Eh, well, you need to know that, eh, they're trying to split up the country and all that. And we don't want that. We, we, we want a, a decent council and, and all.

  84. 15:10

    And that's basically that. That is pretty-

  85. 15:15

    Anyone understand what's going on in Scotland? No? Okay, no Scottish people here. Um, again, we can forward it. We can see, um... And so this is, this is really cool actually that, um, you know, even without fine-tuning the model, it can actually understand specific accents, um, quite well.

  86. 15:37

    So, "You need to know they're trying to split up the country and all that, and we don't want that. We want a decent council and all. And that's basically that."

  87. 15:45

    There you go. I actually now listen to this so often that I can actually hear it. Um, but yeah, you might, you might not be familiar with it. Now, one example that's maybe a bit, um, closer to where you're living, if you're here in the US.

  88. 15:58

    And Mr. President, what you would like to say about the Bangladesh issue? Because we saw, and it is evident, that how the deep state of United State was involved to regime change during the Biden administration.

  89. 16:09

    Then Muhammad Yunus made, uh, junior Soros also. So what is your, your point of view about the Bangladesh issue?

  90. 16:17

    What is the role that the deep state played in, in this situation in Bangladesh?

  91. 16:19

    So what you're seeing here is, uh, president has, um, a translator for English to English, um, translation. So we thought, you know, maybe he can just use this transcription bot.

  92. 16:33

    So if we forward that. Um, again, so even if the audio quality isn't that great, if there's like a lot of background noise, the model is actually very good at kind of identifying that.

  93. 16:43

    Um, yeah, Bangladesh issue, United States. There you go. Um, so yeah. So this is kind of the, you know, one component to it. So we need to, um, you know, understand what the user is saying and then feed that into our LLM, uh, to, you know, have kind of a, a meaningful conversation.

  94. 17:05

    So for the other part, um, you know, we, we don't specifically provide the intelligence layer. So that's where we partner with kind of, you know, the leading, um, model providers.

  95. 17:15

    You can also fine-tune your own model. Um, you know, if you, for example, fine-tune and deploy it on, say, Google Vertex AI. You just need an OpenAI API compatible endpoint, uh, and then you can plug in your custom LLM into this pipeline as well, um, which we can look at in a, in a little bit.

  96. 17:32

    And so now we have the other component. So once the LLM starts streaming out the response, we then want to start streaming the speech as soon as possible. Um, so this pipeline is kind of streaming throughout.

  97. 17:45

    So we have, you know, the, the fastest, kind of snappiest, um, response possible. So for, you know, the actual text-to-speech, what's great with ElevenLabs, you know, we heard we want kind of, um, Brazilian Portuguese accent, right?

  98. 17:59

    So we have a huge library, um, of voices that are available on the platform. So we actually have more than five thousand different voices, um, that you can choose from.

  99. 18:11

    And so if you go into your ElevenLabs account, um, you can go to Voices, and you can explore, um, different voices. Now, if you were sitting, you know, here in this workshop, and you were like, "Oh, I really like this voice," you're in luck.

  100. 18:24

    So you can go to the Voice Library, and you can type in [REDACTED:origin] engineer. I'm originally from Germany. And you find

  101. 18:31

    Me

  102. 18:32

    True success is doing what you are born to do and doing it well.

  103. 18:38

    Does that sound like me? Eh? Okay. It's, it's trained on some of my, uh, YouTube videos, so I think maybe I talk a bit differently, um, in the YouTube videos.

  104. 18:50

    But yeah, this is, this is great. And the thing is-- So I basically clone my voice, uh, and I publish it on the voice library, and so any time you use that voice, uh, I get royalties.

  105. 19:01

    Um, so this is a marketplace. Uh, we actually, uh, recently just surpassed the five million US dollar milestone that we paid out to, um, you know, our voice actors that are kind of publishing their, their voices on the platform.

  106. 19:14

    And so, you know, in this workshop, if you use this voice, uh, I'll be very grateful because then, uh, I can have a coffee later. That's great. Um, no, but obviously, you know, you can use my voice if you want to.

  107. 19:29

    You don't have to. Uh, the great thing is you can set really kind of narrow filters to find sort of the voice that, you know, you want. So for example, if you want to choose, um, Portuguese,

  108. 19:42

    uh, you can just put the language filter to Portuguese, and then you can choose kind of the accent. So here, for example, uh, we want a Brazilian Portuguese. We can then further kind of narrow this down in terms of, you know, sort of gender, age.

  109. 19:54

    So there's certain meta-meta tags that we can apply. Uh, and then we can see maybe here. [foreign language]

  110. 20:04

    Okay, I don't, I don't know. I haven't spent much time in Brazil. Um, so- [foreign language]

  111. 20:12

    Does that sound Brazilian? No? Yeah? Okay. A little bit. Um, so I mean, there's a lot of voices available on the platform. So you can kind of see, um, if you find one that, that sort of fits, you know, the local accent that you're, that you're looking for.

  112. 20:31

    Uh, and then what you can do is you can, um, go and kind of put, you know, all these different pieces together into your conversational, um, AI agent. So here in the dashboard, we can go to Conversational AI, um, and we can actually configure a lot of our agent, you know, right there within the dashboard, and then

  113. 20:50

    we can bring that into our application with, um, the JavaScript SDKs, the Python SDKs, kind of depending on, um, what applications you're building. Um, so we can just do a quick demo maybe of an agent that, um, I had built for, uh, a conference in Singapore.

  114. 21:08

    Um, so, you know, if you're familiar with Singapore, there is, uh, four official government languages. So if you're building applications, you know, for Singapore, you actually need to, you know, provide, uh, English, Mandarin Chinese, Malay, uh, and Tamil.

  115. 21:25

    So these are kind of the official government languages in Singapore. Now, obviously there's other languages, uh, being spoken, you know, Hindi, Japanese, for example, um, you know, as well.

  116. 21:35

    But what you can do is you can, uh, we-- So we currently have within Conversational AI, uh, I think thirty-one, um, different languages that the agent can speak and switch between and identify.

  117. 21:50

    Um, we're working on the next iteration, so the V3 of our multilingual models, which will, uh, bring, I believe, up to ninety-nine different languages. So, um, yeah. You know, no matter where you are in the world, you can kind of get there pretty, pretty good coverage.

  118. 22:06

    Um, so you can see here we can configure the LLM, so we can, you know, choose kind of, uh, any of the, the, you know, big LLM providers that are out there.

  119. 22:15

    We have directly integrated, uh, and kind of co-located with, with our platform. Um, or you can put in your custom LLM, which is just an OpenAI API, uh, compatible endpoint.

  120. 22:28

    And you can put in kind of a knowledge base. Uh, you can upload documents. You can, you know, reference your website. Um, there's RAG built in as well. Uh, depending on kind of the model you're using, um, you might not want to put kind of the entire knowledge base into the context.

  121. 22:43

    So you can, you can en-enable RAG there. Um, and then the tools, you know, this is kind of your standard function calling, um, you know, that you might be familiar with when you're working with, um, large language models.

  122. 22:55

    So depending on the model you choose as well, um, there might be some nuances with how reliably the function calling works. But, you know, kind of the larger ones, GPT-4, um, plus Gemini 2.5, you know, these are generally very reliable with function calling.

  123. 23:12

    And then we have the system tools that are kind of built in. So you can see here the language detection system tool. So we can just enable this. Um, and we have this kind of, you know, configured by default to identify different languages and switch between them.

  124. 23:26

    Um, but you can also override kind of the, the LLM prompt to sort of customize that to your use case. Um, and then maybe we can just give that a little test and see how that works.

  125. 23:38

    Hi there. This is the API Day Singapore Conversational AI Agent. How can I help you today?

  126. 23:44

    Oh, [foreign language] [foreign language] Um, can you tell me what's the name of that ElevenLabs, uh, speaker?

  127. 24:05

    The speaker from ElevenLabs at API Day Singapore is Thorsten Schaeff, who is listed as a developer advocate. Would you like to know more about his session or background?

  128. 24:14

    Yeah, but what, what languages can you speak?

  129. 24:19

    I can communicate in English, Chinese, Hindi, Malay, Japanese, and Tamil. If you prefer to continue in any of these languages, just let me know which one you'd like to use.

  130. 24:27

    Would you like to switch to another language now?

  131. 24:29

    Yes. Um, can we switch to Hindi, please?

  132. 24:35

    Bilkul. Ab hum Hindi mein baat kar sakte hain. Kya aap Thorsten Schaeff ya ElevenLabs ke satr ke baare mein aur jaankari chahte hain?

  133. 24:43

    Um, okay. Unfortunately, I don't speak any Hindi. Um, can you repeat that in English?

  134. 24:53

    Of course, we can continue in English. Would you like more information about Thorsten Schaeff or the ElevenLabs?

  135. 24:59

    So I guess you, you get the gist. Um, so with the language detection tool, there's kind of two, you know, different modes with this. So you can either, you know, as like in the first scenario with my broken Mandarin, uh, I was basically just saying, "Oh, sorry, I don't speak any English.

  136. 25:16

    Can we, like, speak Mandarin?" Um, and it would recognize, oh, okay, you know, even with my terrible Mandarin, it, it was like, "Oh, yeah, he's trying to speak Mandarin.

  137. 25:26

    Okay, so maybe let's reply in, in Mandarin." Uh, so it was doing that. Uh, or we can specifically ask, "Okay, what languages can you speak?" You know, "Can you speak Hindi?"

  138. 25:37

    Uh, "Can we switch into Hindi, please?" Um, so this is kind of the, the built-in language detection system tool that we can use to facilitate kind of these multilingual, uh, conversations, which, which is really nice.

  139. 25:50

    Um, cool. So this is kind of roughly what I wanted to show you, um, sort of as a start. Uh, and now what we can do is kind of we have, you know, the next thirty minutes to play around with this your- yourself.

  140. 26:01

    Uh, and we'll-- We're in the room and, you know, if you have any questions, we, we can answer them. So there's various different ways that you can configure your agents.

  141. 26:11

    So, you know, by default, you can get started in the dashboard, and you can configure kind of a lot of the behavior and functionality, uh, in there. And then, um, you know, if you, if you go back to the resources, we have, um, you know, the documentation.

  142. 26:26

    We have different examples, um, that, that you can use around conversational AI. So for example, there's ex-- you know, we have examples for Next.js to build that into your Next.js applications.

  143. 26:37

    We have examples for Python. Um, if you want to build it, you know, on like hardware devices somewhere, you might want to use Python on like your Raspberry Pi, for example.

  144. 26:47

    Um, so we have the examples there, and you can then, once you've configured that, you can bring that into, you know, your application. Uh, alternatively, you can also configure all nuances of your agents via the API.

  145. 27:00

    Uh, so actually, if you're building a marketplace where you are configuring agents on behalf of someone else, you know, you would generally do that through the API. Um, and we also have, uh, an MCP server.

  146. 27:12

    Uh, so if you're using, um, you know, Claude desktop, you can bring in the MCP server, and you can just tell a natural language, "Oh, please, you know, set up an ElevenLabs Conversational AI agent, um, you know, with this voice," um, and it'll go, and it knows kind of what, you know, API, uh, calls to make to,

  147. 27:33

    to set up your agent. But yes. So, uh, thanks so much for, for joining. Um, if you have any questions, you know, we'll be floating around. Uh, you can as- also ask the questions now, you know, if you wanna ask it kind of in the audience.

  148. 27:45

    But otherwise, um, you know, please just go to elevenlabs.io, um, and create your account if you don't have one already. Um, and then you just go to App, uh, go to Conversational AI, and then you can create, um, here in the Agents interface, you can create a new agent.

  149. 28:05

    Uh, and maybe we'll just start off with, um, kind of the support agent here. Uh, and then we can go through and kind of configure this agent with, you know, your voices, your languages.

  150. 28:16

    Um, and then, yeah, would love to hear a lot of different, uh, agents speak, you know, at, at the end of the next thirty minutes. Awesome. Thanks so much.

  151. 28:25

    Do let us know your questions, and we'll be here for the next thirty minutes to help you set up your agent yourself. Thank you. [audience applauding]

  152. 28:37

    Did anyone have questions that they wanted to ask in the room or... Uh, I think you're, you're welcome. There's like, uh, microphones, yeah, there. Do you wanna just go up to the microphone and ask?

  153. 28:49

    Uh, hello, everyone. First of all, great presentation. Uh, the question is related to you said that it can switch to different languages without fine-tuning it. Like, what is the background process of it?

  154. 29:03

    Can you explain it bit more, like how well it shifts so perfectly that it can interact in the regional languages as well as the accent is also similar to that kind of thing?

  155. 29:14

    Yeah. So here, um, what you can see is in, in my Agent configuration. So I can actually, within the Voices tab, uh, I can assign different voices to the different languages.

  156. 29:29

    So for example, in, um, in, in Singapore, the most commonly spoken Tamil accent in Singapore is the Chennai accent Tamil. Um, and so basically, I went to the voice library, and I found a voice that is C- a Chennai accent Tamil, and I then basically added that to my voice library and just assigned that here.

  157. 29:51

    So in the Voices tab, you can configure the different voices for the different languages. Um, which then means that, you know, the-- so the actual language detection part, that is the automatic speech recognition model.

  158. 30:07

    So the ASR model will actually identify which language, you know, with, um... So it, it basically assigns kind of a, a, a score, like a likelihood score that this is the language that is being spoken, uh, as well as the transcript.

  159. 30:24

    Uh, and so basically we use kind of with the system tool based on the, the confidence score of this language being spoken, um, we then automatically switch to their language k- in the background, uh, and basically use the voice that you configured, um, to, to reply with for this language.

  160. 30:44

    Does that roughly answer the question? Cool. Yeah. Do you mind, um, coming forward so just, uh, because I think it's also being recorded, then-

  161. 30:56

    Yeah, sure

  162. 30:56

    ... then we have it.

  163. 30:58

    Uh, great presentation by the way. Uh, quick question. Uh, other than the

  164. 31:05

    changing of the language or language detection or whatnot, does it have any other ability to make any other actions throughout the call? For example, use case appointment setting.

  165. 31:15

    Mm-hmm.

  166. 31:15

    Does it have actions to maybe call a webhook and to, to, to take a look to see if there's any appointments available either in Make or n8n and then report back, like we can have that in the, the prompts?

  167. 31:27

    Yeah. Correct. So the-- basically, the configuration of that is a, a, um, combination of your system prompt together with the tools. So you can configure, um, custom tools, and so these can be server-side tools, which then would be a webhook call-

  168. 31:47

    Per- per

  169. 31:47

    ... um, you know, to your CRM, to your system. Um, so this is, you know, the, the, the standard kind of tool calling, function calling, um, that the LLM supports.

  170. 31:59

    Um, so you, you know, like GPT-4o or like the, the, the more modern models generally support function calling and structured outputs.

  171. 32:08

    Mm-hmm.

  172. 32:08

    Um, and so as long as the large language model that you're using to power your agent supports function calling, um, you can add your tools and, and this can be a combination of server-side tools.

  173. 32:22

    Um, so for example, you know, as you mentioned, like the scheduling.

  174. 32:25

    Mm-hmm.

  175. 32:25

    Um, you can put in, uh, so for example, we, we also have an example with cal- cal.com. Uh, you can put in the API endpoints for, um, cal.com, and then the agent can actually look up, "Oh, okay, is there availability in the calendar?"

  176. 32:42

    And it can schedule, um... So it can ask for, like the, the email address, and then it, it can schedule, you know, the, the, the meeting, uh, for all the parties kind of through the conversational AI agent.

  177. 32:55

    Okay, one more question.

  178. 32:56

    Mm-hmm.

  179. 32:56

    Uh, what would you suggest, uh, would be a,

  180. 33:01

    a good structure for a conversational agent that has like super low latency? Obviously, the model plays a big role. So obviously, like price and model are like two-- like price and the latency are like the two biggest key factors when we wanna do like a outbound or inbound dialing agent.

  181. 33:18

    What would you suggest for, let's say, an outbound dialing agent or even an inbound to decrease the latency to make it seem more like conversational like? 'Cause obviously, if you use a really good model, it's really large, the latency is just super high, and it just doesn't really...

  182. 33:36

    That- that's not a meaningful conversation to have.

  183. 33:39

    Gotcha. Um, I think that's a good question. I think we have some-- So if you go to the documentation, um, within the Conversational AI, we have some, um, kind of best practices.

  184. 33:53

    I'm not sure if we have like specific-

  185. 33:58

    I do that by testing, but I wanna skip the testing part.

  186. 34:02

    Yeah. Do we have specific guidance on like the best model? It kind of depends on your use case, right? [muffled voice]

  187. 34:14

    Yeah. Um, although I think it's like he's asking about like the best LLM to use, right? For-

  188. 34:22

    Right.

  189. 34:22

    Um, yeah. I think it depends on your use case. Like, you know, depending on kind of the function calling and, you know, how much of that you have, you can probably go down to like, yeah, a Gemini Flash or Flash Lite, um, to like reduce the latency in terms of, um, of that.

  190. 34:44

    But I think on our end in terms of like the voice models, so we use, um... Where are the...

  191. 34:53

    Yeah. We use by default kind of the Flash models for the speech generation. Um, so depending on the languages that you want to support, um,

  192. 35:05

    yeah.

  193. 35:06

    It's just a test. I, I think testing would-- Uh, I think I'll be able to figure out with like proper testing with different models, Flash on, off, all that good stuff.

  194. 35:13

    Thank you so much, man.

  195. 35:14

    Cool. Thanks.

  196. 35:21

    Uh, I have a couple of questions. So number one, what's the total cost per minute coming down to?

  197. 35:28

    You mean like by default? So it kind of depends what, um, pricing tier you're on. Uh, so if you

  198. 35:39

    go to the pricing page, uh, so here Conversational AI. So depending on kind of the, the tier you're on, um, the, you know, there's a certain amount of minutes that are included.

  199. 35:54

    So currently, the pricing is based on call minutes. Um, and then depending on which tier you're on, there is additional minutes that are charged, um, at a specific price.

  200. 36:09

    So it does somewhat depend on, um, kind of the, the pricing tier that you're, that you're on.

  201. 36:20

    If you wanna have like a- an application where you need like really long interaction times, say you wanna do a companion that will, mm, talk to you while you cook, is there something you can do to mitigate that, that cost?

  202. 36:35

    The timing cost? Yeah. That's a good question. I think like for those use cases, it might not be, um, great at the moment because like if they are just running this like twenty-four seven- To, like, be able to talk to someone, um, the cost is, is pretty significant.

  203. 36:53

    Um, that's a good question. I think we-- I mean, there's potentially-- Like, if, if you reach out to the sales team, um, there is some, like, custom pricing that we can do based on the use case.

  204. 37:06

    Um, but I think for now this is charged based on minutes used, like, the, that the session, the session is live.

  205. 37:15

    Okay. And final question. Mm, I tried to do an agent in your dashboard, and it, it had several tasks, so it was one onboarding task and one-- another follow-up task, and it kind of got confused.

  206. 37:29

    Like, it was not, like, identifying when to do one task, when to do the other. So, uh, is there a way to mitigate this or maybe have a multi-agent configuration?

  207. 37:40

    Yeah. So there's, um-- We have what we call, uh, agent-to-agent, uh, agent-to-agent transfers. Um, so one way to do this is to, um, you know, basically set this up.

  208. 37:54

    Um, and so this is, this is a system tool as well, um, where you basically set up different agents for different use cases. So, like, certain use cases you also might use-- want to use a different LLM to power that use case, kind of, you know, depending on, um, which LLM is sort of best for the task.

  209. 38:15

    Uh, and then you can configure different agents for different tasks. Uh, and then y-you can configure kind of an orchestration, uh, setup that basically will then route kind of in the background to a different agent.

  210. 38:29

    Um, if you keep the voice the same, this actually happens somewhat, like, silently without the user actually knowing that they're being transferred. So it's not like-- It's an immediate transfer.

  211. 38:40

    Uh, it just means that you can, you know, sort of develop these agents, um, potentially also across teams where you have, you know, one team that owns kind of this specific agent.

  212. 38:52

    Uh, and then in the background you just kind of switch between the agents, um, for, like, the different tasks.

  213. 38:58

    Great. Thank you.

  214. 39:00

    Thanks. Cool. Any more questions?

  215. 39:06

    Yeah. Um, latency. So in general, like, simple examples usually, you know, great for demos, but when you have something more heavy, more enterprise-grade, and you have more data, more, you know, your RAGs taking time to come back, um,

  216. 39:27

    how do you not have a shitty experience? Because, you know, if we think of it from the end user's perspective, right? They're like, "Oh," and it was like then blank for the next minute or two.

  217. 39:39

    Um, do you suggest fillers? Like, you know, does the conversational AI say, "I'm thinking. Let me think about it." Or h-how, how do you kind of make it more natural?

  218. 39:52

    Because there is gonna be-- Like, in an enterprise setup, right, like, if I have to then

  219. 39:58

    go look up a patient's claim, you know, that might go hit a database. Once that comes back, it goes to other systems, does a whole bunch of things, but it takes time, right?

  220. 40:09

    Yeah.

  221. 40:09

    Like, uh, how, how do you set it up so that, you know, latency that can't be avoided, how do we make the conversational experience better?

  222. 40:20

    Yeah. So there's, there's certain things that you can do. So, you know, generally kind of these are, you know-- If you have like a big knowledge base, for example, like, using RAG is kind of one, um, way of kind of mitigating, you know, that taking too much time.

  223. 40:37

    Um, and then also in terms of your tools, um, when you define your tools, you can, um, configure the--

  224. 40:48

    Where is it? Kind of the response time. So when you, you know, add parameters, um-- Where was the configuration? I think there's a configuration for, yeah, like the timeout.

  225. 41:01

    Um, so basically how long you want to wait, um, for, you know, this tool to sort of come back, um, and also if you want to wait sort of for the response.

  226. 41:12

    Um, so I think the maximum timeout we allow is, like, a hundred twenty seconds. Um, and then

  227. 41:19

    the, the agent will actually, like, say, "Uh, I'm, I'm currently looking that up in the system. Uh, sorry, you know, we're still waiting kinda on the response." Um, so that is sort of built into the tooling here.

  228. 41:32

    Now, you know, depending on your, yeah, like, use case, you probably want to, to, to put that kind of, you know, fairly low because, like, yeah, if you're waiting on the call--

  229. 41:48

    I wonder if you can do something where it's like, um, "Oh, I-I'll call you back." You know? Like, "I'll, I'll take that back and, like, take the action." Or-- But I think for now it would just, yeah, depending on your timeout time-

  230. 42:01

    Okay

  231. 42:01

    ... it basically will wait for the response, and it will tell-- Like, it will stay conversational to like, you know, talk the user through that, "Oh, we're still waiting on, on the tool response"-

  232. 42:12

    Okay

  233. 42:13

    ... um, there.

  234. 42:14

    And are all conversations linear, or can they branch off, come back? Like, while it's looking up something, can it come back in five minutes later, "Oh, you know, I found..."

  235. 42:26

    You know, in the meantime they get other details from the customer or patient or whoever it is.

  236. 42:31

    Um, no. I-- That's a good question. I don't think so. I think at, uh, yeah, currently you would--

  237. 42:43

    I think in that case you would, like, orchestrate it in a way where you put it into a queue. Um, the only thing then is, like, how do you update the agent-

  238. 42:56

    I think like if you're using the-- I, I w- I will need to kind of look this up. There might be a way, like with the WebSockets-

  239. 43:03

    Mm-hmm

  240. 43:03

    ... where, like, as information comes back, uh, you can inject that additional information into the conversation through, um, a WebSocket notification.

  241. 43:13

    Okay.

  242. 43:14

    Um, but yeah, I would need to look up kind of that specific use case, because then what you could do is you put these tasks into a queue, uh, and you work through them in the background, and like as kind of the responses come back, you can then inject them back into the, the context through kind of

  243. 43:33

    WebSocket, uh, events potentially. Uh, but would need to double-check if that's feasible.

  244. 43:40

    Okay. That's all I had. Thanks, man.

  245. 43:42

    Cool. Yeah, thanks. Sweet. Is everyone-- While we're, like, going through the questions, is everyone setting up their agents? Are we...

  246. 43:51

    No? Okay, good. Yeah, I see someone nodding there. Fantastic. Okay. Cheers. Yeah.

  247. 43:56

    Hello. Um, could you explain or demonstrate what happens if you have multiple languages within an input and output, and how that might work if you have different accents for each language?

  248. 44:07

    Yeah. So, so you mean, like, if the person is kind of mixing different languages?

  249. 44:15

    Yeah. Like for instance, if they say, uh, "Explain the difference between schadenfreude and xing sai le ho in Chinese."

  250. 44:23

    Mm. Yeah. It's a good question. I haven't tested that specifically. Um-

  251. 44:30

    Maybe you can try right now. [chuckles]

  252. 44:34

    We, we can, yeah. So what was it? Explain the difference between schadenfreude and?

  253. 44:39

    Xing sai le ho.

  254. 44:41

    Ah, okay. [laughing] Xing sai le ho.

  255. 44:44

    Great. Yeah.

  256. 44:44

    Okay.

  257. 44:47

    Hi there. This is the API Day Singapore Conversational AI agent. How can I help you today?

  258. 44:53

    Uh, can you please explain the difference between schadenfreude and xing sai le ho?

  259. 45:00

    Schadenfreude is a [REDACTED:origin] term that means the pleasure or satisfaction one feels from witnessing another person's misfortune or failure. It is a well-known concept and is often used in English as well.

  260. 45:14

    However, xing sai le ho does not appear to be a recognized term in [REDACTED:origin] or in psychology, and there is no widely accepted meaning for it. It's possible there may be a spelling error or confusion with another term.

  261. 45:28

    Could you clarify what you mean by xing sai le ho or provide more context? That way, I can give you a more accurate comparison.

  262. 45:38

    It might be my pronunciation, but it's, it's meant to be a, a Chinese term.

  263. 45:44

    Thank you for clarifying that you meant a Chinese term. Unfortunately, xing sai le ho does not sound like a recognizable phrase in Mandarin or other major Chinese dialects, at least based on that pronunciation. [laughing]

  264. 45:58

    If you can provide the Chinese characters, a different spelling, or more context about what the term describes or when it is used, I may be able to help you ident-

  265. 46:08

    Yeah, it might be my pronunciation. But yeah, the problem is, like, if you mix,

  266. 46:17

    like, um, you know-- Well, English and [REDACTED:origin] at least are both,

  267. 46:23

    like, Germanic languages, right? But then Mandarin is... So I think there might be, might be challenges there. And like, depending on the more languages you mix, it does get, does get challenging, yeah.

  268. 46:38

    Do you have a recommendation then if you want to build like a language learning application, for example?

  269. 46:48

    Yeah. Um, I wonder if there's certain things you can do with, like, the prompt, the system prompt in terms of, like, improving how it's being picked up. But

  270. 47:06

    yeah, I think, uh, because we're, like, going through text here, um, the, like, the language learning use case is a bit more challenging, especially if you're going, um, you know, Germanic lan-languages versus, um...

  271. 47:23

    Yeah. It's a good question. I, I don't have an immediate answer for you there, but-

  272. 47:31

    Yeah

  273. 47:31

    ... uh... Yeah, actually, you might wanna try, like, a sound token to sound token, like OpenAI real-time.

  274. 47:37

    Yeah.

  275. 47:37

    I wonder if in that case it does better because it can-

  276. 47:41

    Yeah, it does

  277. 47:42

    ... you know, it doesn't go through text.

  278. 47:44

    Yeah. It does that.

  279. 47:45

    So-

  280. 47:45

    Sometimes it, like, switches the accents too, which is kind of annoying 'cause it will try to pronounce Chinese, for instance, in an English accent.

  281. 47:54

    Ah, interesting. Yeah. So yeah, there is, there is challenges with that.

  282. 48:00

    But yeah, that's a, that's a good one. I'll, I'll take that back and see, you know, kind of how we... So I know we have some-- We have a customer in India, Supernova, that does, but it's specifically English learning for, um, the Indian market.

  283. 48:18

    Um, so I think it's a bit of a different use case there.

  284. 48:22

    Do you know if it produces English in a particular accent, or is it, like, using the, you know, Indian phonetic sounds to-

  285. 48:32

    Um, I think there's, there's, like, a case study, uh, ElevenLabs Supernova. So maybe you can, you can look that up. Um,

  286. 48:44

    there's a video, so maybe, yeah.

  287. 48:47

    Yeah. I'll look it up. Thank you so much.

  288. 48:48

    Maybe take a look at that, and then we can-- Yeah, if you, if you connect with us, uh, we can, we can also follow up on kind of some guidance on, on that use case specifically.

  289. 48:59

    Perfect. Thank you.

  290. 49:00

    Cool. Thanks. All right. Hey. Um, I was curious if you're worried about scammers or fraudsters using these tools?

  291. 49:10

    Yeah. So there's definitely, you know, a, a worry with that, like, obviously kinda all this, this technology. So one thing that is kind of very important for us, so you can...

  292. 49:21

    If you go to elevenlabs.io/safety, uh, you can see kind of the safety tools that we're developing, um, you know, uh, in parallel to, to our, um, features. So there's a, there's a bunch of things that we do, um, like specifically, you know, we do, like, life moderation for certain things.

  293. 49:41

    So actually, when you publish your voice to the voice library, you can specify, um, terms that you don't want your voice to say. Uh, and then in this case, we actually have live moderation where, um, we will make sure that your voice isn't used to generate kind of specific terms or, um, sentences.

  294. 50:05

    Um, we also kind of monitor in general, uh, what's being generated on the platform. Um, so kind of the moderation and, um, sort of the, the other toolings that actually with any, um, speech that is generated on our platform, we mark watermark it, uh, actually to the extent that we can trace back which account generated, um,

  295. 50:30

    this specific speech. Uh, so if we identify, um, fraudulent activity, we can actually trace back which account generated kind of that, um, and, you know, can kind of ban them or, you know, provide, um, kind of information to the authorities, um, as needed.

  296. 50:50

    Yeah. And so the other things is just kinda in terms of, um, for, uh, we have... Like, where was the... We have, like, the voice capture, um, that we developed.

  297. 51:03

    When you are creating a, um, a professional voice clone,

  298. 51:09

    uh, we actually generate kind of a random sentence that you need to read out to verify that, you know, you have permission to clone this, this voice. Awesome. Um, so yeah, we, we do, you know, with kind of all the technology that we develop, um, we do put, uh, quite a large amount of, you know, focus and

  299. 51:31

    effort into, uh, safety tooling. Um, but yeah, there is obviously always a concern that your technology is being used for fraudulent activity. But I think so far, um, you know, we've been trying to mitigate that with, like, the safety tooling.

  300. 51:48

    Definitely, yeah. Looks like a lot of good guardrails in place. Thanks. Thanks.

  301. 51:58

    Yay, you're back. [laughs]

  302. 52:00

    Yeah. You wanna go... So just building on her one, like, um, in our place, um, if a patient is asked, "Hey, you know, how are you feeling about this?"

  303. 52:11

    And say they are, um... They, they try to speak in English, they might hold a part of the conversation in English, and then they might jump to Spanish, uh, Spanish and Portuguese, and come back to English for, you know, like, when they have to describe something that they can't in English, uh, they kind of jump back

  304. 52:36

    to, you know, the, the language that they are most comfortable with.

  305. 52:39

    Mm-hmm.

  306. 52:40

    Sometimes they jump around different languages. Uh, how would-- Because, like, the way you explained it, it kind of, you have some kind of a router that checks what kind of language it is and then shoots it off.

  307. 52:56

    But within a conversation they, they kind of jump between, like, you know, they'll explain a few things and then they'll go a few words with a very, you know-

  308. 53:06

    Yeah

  309. 53:06

    ... Portuguese words and stuff.

  310. 53:08

    So, like, for you, this, this is like English, Spanish, Portuguese, kinda all mixed-

  311. 53:14

    Yes

  312. 53:14

    ... together?

  313. 53:14

    Sometimes. Like, if you just ask them how you're feeling, okay, then they come back, "Hey, you know, there's this, you know, I took this medication, it hurts," you know?

  314. 53:22

    And then if you say, "Okay, where is it hurting? How..." Then they suddenly kind of, you know, they re- regress to whatever language is most comfortable to them to explain their thing.

  315. 53:37

    Gotcha. Yeah. I mean, yeah, you can see here that, like, the transcript, it actually correctly identified, you know, schadenfreude, um, because technically it's also an English word, right? But then, like, on the, on the Chinese word, it just completely, you know-

  316. 53:53

    Oh, yeah.

  317. 53:53

    Well, you know, you can blame partly my pronunciation. Probably you can blame it a lot. But, um, yeah, I wonder if, like,

  318. 54:04

    a native speaker... Yeah, I don't, I don't have exact benchmarks on, like, you know, how much, like, the transcript gets worse the more languages you introduce kinda in the same.

  319. 54:17

    So I think, like, generally if you have two languages intermixed, it tends to perform okay. But, like, if you're, like, now having, like, three different languages, um, it just, you know, progressively tends to, to get worse.

  320. 54:35

    Okay.

  321. 54:35

    But I don't have, um, exact benchmarks on, like-

  322. 54:39

    No, no

  323. 54:39

    ... you know, how many languages sort of... Yeah.

  324. 54:42

    Okay, cool.

  325. 54:43

    So-

  326. 54:43

    No, that's, that, that's good to know. That's what I came up with.

  327. 54:46

    But yeah, it, it would be worthwhile if you have, like, recordings of, like, some of that to, like, put it through our, um, transcription model and see kind of how, how it performs in, like, identifying that.

  328. 55:01

    That, that would be interesting, yeah Cool. Thank you. [clears throat]

  329. 55:06

    Uh, save this question toward the end 'cause it's kind of non-related. So, I worked on a project where we used ElevenLabs, uh, for, uh, the voice track of our avatar.

  330. 55:15

    Okay.

  331. 55:15

    Uh, and ElevenLabs functioned well, but we had a lot of more downstream issues in terms of, like, lip sync and, like, uh... I think someone mentioned, like, slugs and timing and other things.

  332. 55:25

    So is there any plan for ElevenLabs to come, like, I guess further down the stack in terms of, like, avatars? Or is, is that even something you're thinking about?

  333. 55:35

    Uh, interesting. So you, you... Did you build, like, the, the lip syncing model and, like, avatar stuff on your-

  334. 55:43

    Uh, no

  335. 55:43

    ... into yourself?

  336. 55:44

    So, uh, we, like, within the, like, NVIDIA Tokyo stack, so they have, like, a, a stack, and they have their, like, Riva voice model, and we kinda switched that out for ElevenLabs.

  337. 55:54

    Uh, so they have, like, the full stack of, like, the avatar, and then ElevenLabs is just the voice port- portion of it. Yeah.

  338. 56:00

    Ah, okay. And, uh, sorry, which, which stack was that? The NVIDIA-

  339. 56:06

    Oh, NVIDIA Tokyo. Yeah. It's like-

  340. 56:08

    Oh.

  341. 56:09

    Yeah. It'll go from, uh... Like, we did everything with the, the GPU, but then they have, like, the visualization-

  342. 56:16

    Okay

  343. 56:16

    ... and, uh, you can just plug your voice model or, or in. Uh, it's Tokyo, like T-O-K-K-I-O.

  344. 56:24

    T-O-K-K... Ah, I see. I'm, I'm thinking about Japan. Is it this one?

  345. 56:30

    Yeah.

  346. 56:32

    Oh, interesting. Okay. Yeah. I, I personally don't have,

  347. 56:39

    um, experience with that one. So I know that we're mostly working with partners like Hidra and HeyGen kind of for sort of the, the avatar-

  348. 56:50

    Mm

  349. 56:50

    ... side of things. Um, I don't know. Paul, do you know any? No. So this is something I, I would need to come back to you and, like, look into.

  350. 57:01

    Um, it's, it's interesting. So, like, you're saying out of the box it uses, like, an NVIDIA model for speech generation or-

  351. 57:09

    Yeah, yeah. But, uh, our client, which I assume is one of your partners, we can't talk about that, but, like, our client couldn't use the NVIDIA model.

  352. 57:17

    Gotcha.

  353. 57:18

    And had a contract to use ElevenLabs model. So-

  354. 57:20

    Okay

  355. 57:21

    ... it's kinda-

  356. 57:22

    Interesting. Yeah. Sorry, I don't ha- I don't have a good answer for you there-

  357. 57:26

    Mm-hmm

  358. 57:26

    ... right now. But yeah, this is interesting. We c- we can go back to the team and see, um, if there's any resources that we can, we can give you in terms of, like, improving that.

  359. 57:38

    No worries. Thanks. Appreciate it.

  360. 57:38

    But interesting use case. Thank you. All right. Last question.

  361. 57:44

    Last question. Nice. Uh, hi, Thorsten. Thanks for the presentation. Um-

  362. 57:48

    Thank you.

  363. 57:49

    I have a question regarding the transcription model around adding custom vocabulary. Like, uh, at the company I work for, we use a lot of three-letter acronyms. And, uh, let's say, let's say I want to have the model read out SAP as SAP and not SAP.

  364. 58:05

    Is there a way to tell it to do that? And is there a way to, uh, like, tell it to read words a certain way and kind of nudge the interpretation of what I say towards certain words that we use in our vocabulary?

  365. 58:19

    Interesting. Yes. So you have this both... So you have this use case both on, like, the s- the speech-to-text, so you need to correctly identify the acronyms. But then also you need the agent to reply back with the correct...

  366. 58:34

    So, like, for the reply back, we do have, um... Have you seen the, like, um, pronunciation dictionaries? Um, so we have a way for you to provide, um, you know, pronunciation dictionaries with, like, kinda phoneme alphabets, uh, to actually, you know, identify specific, uh, you know, basically

  367. 58:59

    like here, tomato. Tomato, I guess. Tomato. Tomato. Um, and so you can provide

  368. 59:09

    for, for that, you can provide the pronunciation dictionary, dictionaries to make sure the text-to-speech pronounces, you know, the acronyms and the words in the way that you want them to.

  369. 59:21

    Now, the other side of, like, the speech-to-text, that's an interesting case. I don't think we have,

  370. 59:33

    uh, a way to, like, fine-tune that specifically for different acronyms. That's a good question. So

  371. 59:48

    you know anything there? No, right?

  372. 59:51

    You can do a normalization layer when you

  373. 59:56

    talk to the LLM. So you put in a prompt which-

  374. 59:57

    Oh, interesting. Yeah. So you, you can-- to a certain extent, you can do it through the system prompt, where you put in kind of a normalization layer to basically identify things in the transcript that are acronyms, and then basically have the LLM sort of massage that into, to what you want.

  375. 1:00:18

    I think that's what you were saying, right? Yeah.

  376. 1:00:20

    Mm-hmm.

  377. 1:00:21

    So that could be interesting to see if that works well.

  378. 1:00:24

    Yeah. Great. Thank you.

  379. 1:00:25

    Have you, have you tried it out already or...?

  380. 1:00:27

    Um, where I was coming from is, uh, the company is using, um, like in their own custom chatbot called Joule, and it's like the unit of work. But whenever I read transcripts, it's oftentimes used as jewel as the diamond.

  381. 1:00:40

    Gotcha.

  382. 1:00:41

    And so that's kind of the struggle that I'm facing.

  383. 1:00:43

    Okay. Is this actually at SAP?

  384. 1:00:46

    Mm-hmm.

  385. 1:00:46

    Nice. I, I'm an SAP child myself.

  386. 1:00:50

    Mm-hmm.

  387. 1:00:50

    My, my father was early S... Well, okay. Anyway, too, too much information. Cool. Uh, yeah. Thanks, thanks for that. We'll, we can, we can chat some more and see sort of if that's something we, we can get going.

  388. 1:01:04

    Sweet. Uh, yeah. And with that, we're at time. Um, yeah. Thanks again. Thanks so much for joining. Please do, um, you know, connect, uh, find the resources, fill in the form for the credits.

  389. 1:01:18

    Uh, yeah. I'll, I'll leave this up, uh, in case you haven't had a chance to scan it. But yeah. Thanks so much for joining. Enjoy the conference and, uh, we will also have a booth at the expo, so if you come up with some more questions, you can come, uh, find us there.

  390. 1:01:32

    Thank you. Danke schön. [laughs] [outro music]