AI Engineer Europe 2026
Beyond Transcription: Building Voice AI That Actually Understands Conversations
About this talk
pyannoteAI co-founder and chief science officer Hervé Bredin explains why speech-to-text alone cannot capture who said what or how conversational context changes meaning. He introduces pyannote speaker diarization, its integration with Whisper and downstream language-model workflows, applications in meeting summaries and podcast analysis, and diarization error rate. A live Python demonstration uses pretrained Hugging Face models with PyTorch and Apple MPS to examine speaker attribution, interruption, and overlapping speech in a recorded telephone conversation.
Chapters
- 0:00Introduction: Hervé Bredin, pyannoteAI, and speaker diarization
- 2:12Why transcription needs speaker attribution and conversational context
- 7:22Speaker diarization, speaker identity, and speech detection
- 12:27DER and live Python diarization demonstration
- 20:15Overlapping speech, interruptions, and closing
Talk transcript
- 0:00
[upbeat music] Uh, good morning, everyone.
- 0:16
Thanks for, uh, being here to the, uh, Voice and, uh, Vision session. Uh, so I'm, uh, Hervé Bredin, chief science officer and co-founder at pyannoteAI. Uh, so I'm going to talk to you, uh, today about, um, conversations, understanding conversations and, uh, uh, what you can do on top of, uh, transcription.
- 0:37
Uh, so a quick words about myself. So I've been a, an academic researcher all my life until, uh, two years ago when I, uh, started this company. Uh, oops, sorry.
- 0:48
Uh, what happened? Yeah. Uh, so basically, yeah, I, I, I worked on this topic called speaker diarization, which I'll, uh, introduce a bit later for those of you who don't know this, uh, weird, uh, word that is tricky for me to pronounce.
- 1:03
Uh, but basically, over the years, I built an open source toolkit called, uh, pyannote, uh, which focuses on speaker diarization, and that became quite popular over the years, in particular since, uh, uh, OpenAI, uh, uh, released, uh, Whisper speech-to-text, uh, open source models.
- 1:19
Basically, Whisper was, uh, some kind of revolution, uh, in terms of STT. Uh, the fact that it was free, that it was, uh, very good and, uh, but it didn't provided the actual, uh...
- 1:28
didn't provide the actual names and tags of the speaker, so people naturally turn to pyannote to combine the two. And you, you can see that on the, uh, inflection points on the GitHub star history.
- 1:39
And, uh, just a few words, I happen to turn [REDACTED:age] uh, uh, in just one week from now, and we are almost at 10K, uh, stars on the GitHub.
- 1:48
So please, uh, give me a, a birthday gift by, uh, just, uh, going to the website and, uh, to the GitHub website and add your star. That will make my, uh, birthday.
- 1:57
Um, so, um, let, let, let's go to the, to the core of the presentation. So you all know what transcription or a speech-to-text is. Basically, that's answering the question, what was said from the recording or streaming of a, a conversation.
- 2:12
Basically, you get from the audio, and you get a sequence of words, uh, uh, as output. Uh, that's, for instance, what Whisper does. Uh, but, uh, without the actual, uh, names of the speaker, the usually the conversations on the transcript are not really, uh, understandable.
- 2:29
So the next step after transcription is basically to attribute, uh, a, a speaker tag to each, uh, of the words. So that's, uh, I call here a speaker attributed transcription.
- 2:39
And so that, that answers the question, who said what? But for some application, that's enough. Like for instance, uh, uh, for maybe for meeting note takers to make a summary of a, of, of a meeting, that might be enough to assign an action point to the right people to, to know exactly that, uh, Hervé said that and,
- 2:57
uh, John said, said th- this other thing. But we-- to understand the conversation, uh, a bit better, uh, we, we might want, uh, uh, to go, uh, slightly, uh, better.
- 3:09
Uh, sorry, I forgot about this slide. So basically, uh, there are many cases where, um, knowing who said what is as important as what was said, actually. So I have here a few examples.
- 3:19
Like, for instance, uh, video dubbing, automatic video dubbing. Um, uh, s- knowing who said what is actually really important to put the actual, uh, uh, good correct voice to the, uh, right speaker.
- 3:33
So i- if you want to automatically translate, uh, um, a video from one language to another, you want the voices to be consistent as well. So for that, you need, uh, to know who's, uh, who, who spoke when.
- 3:46
Same for meeting note takers, and there are other applications like, uh, for instance, podcast, uh, what I call podcast intelligence, like being able to track a speaker across different, uh, uh, episodes of a podcast or, uh, a- across different podcasts, finding, uh, the same guest across, uh, multiple podcasts.
- 4:05
But yeah, so, uh, I was a bit, uh, o- over, uh, just before. But, uh, the, just knowing who said what is sometimes not enough to really understand the conversation, the conversation, how it goes.
- 4:16
So knowing who said what and when actually, uh, brings even more information. For instance, in this example, um, uh, without knowing exactly when each words were pronounced, you, you can't detect that, uh, actually, the black speaker interrupted, uh, the green one.
- 4:32
You can't really detect that, uh, maybe this small backchannel that, uh, the, the, the green, the, sorry, the black speaker, um, uh, does during the speech of the, uh, of the green speaker.
- 4:43
If you miss it, basically, you, you don't really understand that, uh, maybe the black speaker is actually agreeing, nodding, uh, at what the, the green speaker says. So knowing exactly when this small word, uh, has been pronounced, uh, is very important.
- 4:58
And same, without precise timestamps, you can't really know whether I made a... I make a, a pause in between two, uh, speech terms, which might, uh, convey additional information about, uh, my state of mind, my, I don't know, the, the, the point that I'm trying to, to make in the conversation.
- 5:14
And we can go even a step further. Like we, we might want to know who said what and when and how they said it. Like, uh, I have a few examples here.
- 5:24
If, uh, some of you are laughing at some point, is it because, I don't know, I said something stupid or I was actually funny? Uh, same for coughing. Uh, maybe there's something in the conversation that makes me, uh, uncomfortable and, uh, I'm coughing for some reason to, uh, to hide be- behind that.
- 5:41
And, um, yeah, so I also mentioned stress, disfluency, prosody. Um, the fact that, uh, I, uh, stress on particular words in a sentence might actually change completely the meaning of the sentence.
- 5:54
Uh, I have this example in mind, like, uh, I don't know, in the sentence, "The dog, uh, ate, uh, the cake." If, uh, I don't know, you're talking to a child and, uh, the, the child tells you, "The dog ate the cake," or, "The dog ate the cake," or, "The dog ate the cake," uh, the, the meaning
- 6:11
might be slightly different. So, uh, stressing particular words, uh, being able to have a, a voice AI actually understand the stress and, uh, and, um, this kind of, of, um, low-level details might bring some more, uh, information to the,
- 6:26
to any downstream LLMs or, or whatever tools that you use after this first, uh, enriched transcription step.
- 6:35
And if we take even a step, uh, a step back, uh, um, knowing who's, uh, talking to who. Uh, am I, uh, here in this example, I'm probably can be seen as this, uh, gray speaker addressing to all of you.
- 6:49
But at some point, if, uh, uh, at the end of this talk, I, uh, you do have questions, then, uh, I will be probably the, the green one talking to one particular person answering the particular question.
- 6:58
So knowing that, having a higher level view of the conversation, it also opens up a, a, a, a new bunch of, um, applications.
- 7:09
And all of this, uh, also happens in a acoustic environment that gives you information about the context, whether it's in this quiet room, uh, for now, uh, or if I'm in the street.
- 7:22
Uh, this bring more context, uh, to, to, to the game, and you can actually, uh, take a, a better decision about, uh, what's going on, why a person is actually saying what, what they are saying and, and et cetera.
- 7:34
So all of this is basically the, what we are trying to work on at, uh, pyannoteAI, but we, uh, started with a, a smaller problem called speaker diarization, and, uh, that's really what I want to talk a-about, uh, today.
- 7:50
So speaker diarization is the... I, I... Yeah, i-is, is usually... Yeah, what I, what I wanted to say about this slide is that, uh, who said what is, uh, usually just as important as what was said, and you can tell by instance by going through, uh, to the Hugging Face, uh, uh, model repository and filter, um, all
- 8:09
the models there, uh, with the ones that have the audio tag. Uh, where is my mouse? Yeah, here. So this is what I did, and then you sort them by, uh, by number of downloads.
- 8:19
And basically, uh, the... among the, the top, uh, here, seven models that, uh, are, are at the top of this list, three of them, the, the, the first three are actually related to speaker identity and speaker diarization, and the, the, the other ones are STT.
- 8:33
So that, that's one way of showing that, uh, speaker diarization is actually quite, uh, quite important on, on, i-in the community.
- 8:42
So, uh, going, uh, into a bit more detail about what speaker diarization is. So, as I said, it's answering the question who speaks when. So starting from the recording of a conversation, uh, basically, the first step that people usually do is, uh, basic voice activity detection.
- 8:59
So you basically tell whether someone is speaking at, at one time or not. Anyone is speaking at, at one time or not.
- 9:06
And then you can go a step further. Uh, you can actually segment those speech regions into smaller speech turns. Um, so thi-this is what I call segmentation, and it, it includes finding, uh, speaker change points between, uh, i-into those long, uh, in those long speech regions, you might find a, you know, a speaker change point, and, uh,
- 9:26
also find regions where, uh, people are actually interrupting each other. Like, uh, in, uh, in the first example here, there's definitely here someone interrupting, uh, someone else, and it's possibly here some kind of back channel where someone is actually, uh, uh, nodding or, uh, saying, "Mm-hmm, okay," th-this kind of stuff.
- 9:45
And this is the kind of small speech turn that, uh, you don't want to miss because sometimes they actually convey, uh, the most important information of the conversation. If it's a small yes here, uh, and if you miss it, uh, then, uh, you, you have no idea what's the, the, the state of mind of this speaker.
- 10:05
But that's not yet speaker diarization. Speaker diarization goes all the way, uh, to actually assigning a speaker identity to, uh, each, uh, speech turn.
- 10:16
And, um, so here in this example, the, the speaker diarization system automatically detected that there are two speakers, the green one and the, uh, black one. Um, but, uh, we usually don't have any prior knowledge on the actual number of speakers in the conversation.
- 10:33
You might have some kind of, uh, uh, guidance on, uh, I don't know, if, if you are a meeting note-taker, for instance, uh, you, you have the list of attendees.
- 10:41
Uh, so you might have an idea of the number of people that were invited, but that doesn't mean that, uh, someone, um, two people join from the same channel or, or, uh, that an attendee that was not invited actually, uh, join in finally and, uh...
- 10:56
So, uh, that, that makes the problem, uh, kind of difficult in the sense that we don't know in advance the number of classes that we're supposed to detect, uh, contrary to other, uh, let's say, classical machine learning problems.
- 11:08
And also, we don't know the identity of the speaker, uh, that we're supposed to, uh, uh, diarize, in the sense that speaker diarization does not really output, uh, John, Hervé and, uh, and, uh, and Jack.
- 11:23
Uh, but basically, it's more like speaker one, speaker two, speaker three, and they can be permutated, and it's still correct. For example, i-in this example, if I just change, uh, green into black and black into green, the diarization is still exactly the same, uh, from the point of view of the evaluation metrics that I'm going to talk
- 11:40
about and, uh, uh, from the point of view of the actual, uh, speaker diarization task.
- 11:46
So that makes the, uh, this problem, uh, kind of, uh, difficult, and that's why even though, uh, the community, uh, has been working on this topic for a long time now, it's still not solved.
- 11:56
Uh, there are other reasons, like the fact that we have to detect overlapping speech, uh, to handle overlapping speech, uh, that we have to, uh, uh, take into account very short speech turns, that don't have to take into account- Imbalance between, uh, the speech time of multiple speakers in the, in the conversation, all of that makes the
- 12:14
problem difficult. On top of the usual, uh, acoustic conditions, uh, problems and everything that any speech processing, uh, um, um, task, uh, have to deal with.
- 12:27
So, uh, I'm going to switch to a demo to, to show you, uh, a bit more how the, uh, diarization system I evaluated. Um, maybe when you, you, you stumble in- into benchmark, you, you find this, uh, DER acronym, which is, uh, diariz- stands for diarization error rate.
- 12:42
And, uh, so hopefully, um, this will work. So, uh, I have here, uh, a Python notebook, uh, that, uh, I've prepared for, for this talk. It's actually already, uh, available on GitHub.
- 12:56
I'll show you the, the QR code at the end of the presentation, so you can play with it. But so the... in this example, I have a conversation between two [REDACTED:gender], uh, talking over the phone.
- 13:05
Hello?
- 13:06
Hello.
- 13:07
Oh, hello. I didn't know you were there. [laughs]
- 13:08
Neither did I.
- 13:09
Okay. Wait, I thought, you know, I heard a beep. This is Diane in New Jersey. Um-
- 13:13
And I'm Sheila in Texas, originally from Chicago. [laughs]
- 13:16
Oh, I'm originally from Chicago also. I'm in New Jersey now, though. [laughs] Okay. So you, you get the idea, a 30-seconds conversation between two [REDACTED:gender]. So this is the expected output of a perfect speaker-attributed, uh, transcription.
- 13:27
This was manually labeled. And now, uh, what I'm going to, to, to do this here is, is actually run a piano open source, uh, Community-1 model on, on the same file.
- 13:37
So it's actually, uh, running, uh, right now on my Mac. So that's why here I'm using this, uh, MPS, uh, PyTorch, uh, backend. So how it works is that, uh, you first download the model from, uh, Hugging Face with this, uh, line of code, then simply, uh, uh, apply it on, on the audio file.
- 13:54
So now, uh, it's finished and the... there's the, this prediction variable that contains, um,
- 14:01
uh... Sorry, I should have done that on the other tab because this was the one that was already processed. I wanted to do it live, so let's do it live now.
- 14:10
So it's running. And so the, the next step is to actually visualize the errors that, uh, this system, uh, made. So at, at the top here, you have the reference annotation, so the expected output of diarization.
- 14:22
And here is the output of, uh, Community-1, uh, speaker diarization, uh, pipeline, uh, that is open source and free to use. So it makes three kinds of mistake, as I was saying.
- 14:32
Uh, like a confusion. Here, it got the speaker wrong. Uh, it can also make, uh, false alarms. Uh, so false alarm is when it detects speech when there's actually no speech in the ground truth.
- 14:43
And, uh, misdetection is the other way around. And misdetection can actually, uh, happen, for instance, here during overlapping speech, when we actually detect only one of the two speakers here.
- 14:55
And then once we have these, uh, uh, basically errors of, uh, false alarm, confusion and, and misdetection, we, we can actually compute the diarization error rate. Uh, so this is with this library called pyannote metrics.
- 15:08
And, uh, basically, it gives us, in this example, a five percent diarization error rate, which is basically the sum of the confusion, uh, uh, false alarm and misdetection divided by the, the total, uh, duration of, uh, of speech in, in this, uh, in this file.
- 15:25
And, um, so w- we have this other model that I'm currently running. Hopefully it work, it, it works. So it's running, uh, right now on our, uh, cloud API.
- 15:36
Uh, basically, it's, it's, uh, uh, a better model, that Community-1, that we call Precision-2, that is, uh, uh, running here. Hopefully it will work, but as for every demo, it might fail.
- 15:48
If... No, it worked. And, uh, this is the kind of output that you can get. So you see that, uh, it got the, the s- the speaker right here and, uh, makes, uh, slightly less mistakes in terms of mi- uh, false alarm and misdetection.
- 16:00
And overall, in this example, I think we got a, a three percent, uh, diarization error rate here.
- 16:07
So let me switch back to the, to the presentation. Uh, so this was about, uh, benchmarking, uh, speaker diarization model, and I'm often asked how well the state-of-the-art speaker diarization works today.
- 16:19
And it- it's a difficult question to answer because really, depending on the use case, it might, uh, vary completely. For instance, in this example, we have a here conversation telephone speech like the one we just, uh, listened to, like two person over the phone, and we, we can go down to eight percent, uh, diarization error rate.
- 16:35
The best system, uh, does that. But if, if you are now in a restaurant with many friends, with lots of background noise, uh, even the best system reaches like forty-one percent diarization error rate.
- 16:46
So it's far from being a solved problem, uh, but, uh, we, we are working on that.
- 16:52
And, um, so once you have speaker diarization and you have transcription, shouldn't be that hard to actually find a speaker-attributed transcription, right? It's just a matter of, uh, you know, assigning a word to a speaker, a speaker to a word.
- 17:06
But actually, that's not that, uh, easy. And I'll, um... There, there are many reason why that's not that easy. The first reason is that most, uh, speech-to-text models are actually trained on single speaker data.
- 17:19
As soon as you apply them on multi-speaker data with overlap, with speaker change, with all kind of mess, they completely, uh, uh, they fail miserably. So for instance, when you look at the open ASR leaderboard from, uh, Hugging Face, if we look, for instance, at NVIDIA Parakeet that we are using, uh, that, that I will be using
- 17:36
in a, in a demo right after that, they report eleven point four, uh, percent word error rate.
- 17:43
When we apply the very same model on the same AMI, um, on, um, on our side, we, we get twenty-six percent. So the question is, w- why is there a difference?
- 17:55
Actually, it's in the way the, those benchmark are set up here on the, uh, uh... For, for example, in this AMI dataset, that's, uh, uh, a meet, um, a datasets with meetings, uh, between four to five people, I think, in meeting rooms.
- 18:09
And they are basically microphones like the one I'm, I'm, uh, having here, so a headset microphone, as well as one microphone in the middle of the, the table. And, and, uh, those numbers here, uh, on the OpenAI ASR leaderboard are based on the headset microphone.
- 18:24
Those numbers here, uh, are based on the microphone that is in the middle. And so on one side, you have a single speaker speech. On the other side, you have multi-speaker and even distant microphone speech.
- 18:34
So that's why it degrades, uh, a lot. And so that, that, that's... that was my point of the, of this slide. So, uh, when doing speaker attributed transcription, the, um,
- 18:44
the reason why it, it might go wrong is either because STT doesn't work great and, um, uh, it's usually the case that, uh, they don't generalize, uh, very well to multi-speaker recordings for, as I was saying, distant microphone, uh, speaker change, crosstalk, interruptions, you name it.
- 19:02
Uh, maybe code switching as well when you change language in the middle of a sentence.
- 19:07
But also it's, it might be because of the actual reconciliation between diarization and, uh, STT timestamps. Though it may sound, uh, uh, obvious how to do that, as I was saying, I will hopefully show in this, uh, live demo once again that, uh, it's, it's not such a, an easy problem because, uh, STT does not transcribe overlapping
- 19:28
speech well, uh, because, uh, the timestamps disagree between STT and, uh, diarization, and because sometimes diarization will detect speech that the transcription will not transcribe and the other way around.
- 19:43
So, um, second part of the demo, uh, about speaker attributed transcription. So, uh, first, I'm going to, uh, apply, uh, this Parakit model, uh, from NVIDIA that does, uh, by the way, a great job at, uh, transcribing, uh,
- 19:58
uh, uh, th-this sample. And so, uh, Parakit gives us this kind of output. So we have this sequence of words, and for each word, we actually have, uh, the, the corresponding timestamps.
- 20:11
Okay? So, uh, I, I can play it quickly. Hello?
- 20:15
Oh, hello. I didn't know you were there. [laughs]
- 20:16
Neither did I.
- 20:17
Okay. I thought, you know, I heard a beep. This is Diane in New Jersey, um-
- 20:21
Well, you, you get the idea. And then, um, the question is, um,
- 20:28
now we have our best diarization. This is Precision-2 output here. This is Parakit output here. And the question is, uh, now we need to assign a, a speaker to each word.
- 20:37
So le-le-let's do it, uh, step by step, for instance. Uh, so in this example, even though the timestamps of the, the words, uh, are a, a bit shifted, uh, on the left here, it's obvious that, uh, it's probably the yellow speaker who spoke the, the gray, uh, word.
- 20:53
But if you go just the third word, so now we have this. So, uh, we have this, uh, word, uh, the word O here, uh, that is, uh, in between those two, uh, speech turns according to diarization.
- 21:05
Which one do you assign it to? Um, so we, we... I, I can listen to it. Oh, hello.
- 21:11
Oh, hello. Right. It's not, it's not quite, quite sure. And there are many other, uh, positions where it happens here same. There's a word here, okay, that, uh, apparently according to diarization, there are two speakers speaking and, uh, only, only one word from the, from the transcription.
- 21:26
So that, that's why, uh, at, uh, pyannoteAI, we came up with this, uh, new, uh, STT orchestration thing that basically does, uh, its magic and, and does the job for you.
- 21:35
So here I j- I just submitted, um, to our cloud, uh, API, um,
- 21:42
uh, a job to transcribe and, and, and use both Precision-2 diarization model and Parakit transcription model and take care of the... what I call this, uh, reconciliation between the two.
- 21:54
And you end up with this kind of, uh, with this kind of output. Uh, what's nice is that so now you... it did the, the, the job for you, but in particular, I, I like to go to this particular example.
- 22:06
Even where there was overlapping speech before, it actually managed to, uh, interleave the, the words from those speaker. Let me play th- just this part and focus on this, uh, area maybe.
- 22:16
In New Jersey, um-
- 22:17
And I'm Sheila in tech- And there, there's actually a and, uh, that, uh, or this is the um, here that, that is, uh, actually, uh, uh, the, the, the, the, the, the yellow [REDACTED:gender] a- actually, uh, has this um, at the end of their, her speech and, uh, the, the blue one actually interrupts her.
- 22:37
But they, they, they're actually overlapping, and we managed to get them right.
- 22:42
Um, yeah, so th-this is, uh, the end of my talk. I, I, I wanted to, uh, give you a f- a, a few minutes. I have two minutes left for questions, but I wa- just wanted to say that this demo, uh, that I, I did is already available on, uh, in the pyannoteAI/tutorials, uh, GitHub repo.
- 22:58
So in there, you, you, you'll be able to play with pyannote.audio, open source toolkit for diarization, with pyannote-metrics for evaluation, uh, with iPyannote for this nice, uh, visualization, uh, widget that we've, uh, implemented.
- 23:10
I stands for, uh, not iOS, iMac or whatever, but f- really for interactive and, uh, also the SDK to, to play with our premium models. And, uh, I'll stop here, and, um, happy to take, uh, a few questions.
- 23:23
We have one more minute left. [clapping] Yes, please.
- 23:34
Um, so what's the trick, uh, that you use to resolve the conflicts?
- 23:40
So-
- 23:40
Is it part of the model training, or is it some sort of like heuristics that you have around to identify who to attribute the word to?
- 23:48
Yeah. So the question is, uh, what's the trick to, uh, actually solve this problem of a reconciliation? Uh, so that, that's a, a, a proprietary trick. But what I can, what I can say is that we, uh...
- 23:59
there's already part of this trick that is available in the Community-1 model, which we call exclusive diarization. Basically, uh, what we do is that we find a way when there is overlap to actually select, uh, the, the most, uh, likely of the two speaker that will be transcribed by the STT model.
- 24:14
So that simplifies actually the reconciliation between the two. So that, that's one major part of the, uh, of, of the approach.
- 24:21
Part of the model training is like a se- couple of heuristics you have to-
- 24:25
Yeah. So, so the question is about, uh... I, I'm repeating because I've been told I need to repeat questions. Uh, so it's not part of the model training in the sense that we really plan to support any kind of STT without having to change the STT model itself.
- 24:38
So really, it's supposed to work with any STT, uh, even fine-tuned ones that you might have internally for... because you, you fine-tuned them for your particular use case and nobody has it.
- 24:48
Uh, you can combine it, uh, like that.
- 24:50
Yes.
- 24:52
All right. Yeah, eighteen seconds, so I guess we need to stop. I'm sorry. Maybe we can talk offline. Otherwise, uh, they, they'll, they'll beat me, I guess. [laughs]
- 25:00
Thank you very much. I'm sorry. [clapping] [outro music]