AI Engineer Europe 2026
Beyond Transcription: Building Voice AI That Actually Understands Conversations
Read the talk
Beyond Transcription: Building Voice AI That Understands Conversations
Words alone miss interruptions, backchannels and speaker changes. Hervé Bredin shows how diarization works, how to evaluate it, and why joining its output to a transcript remains difficult.
From a talk by Hervé Bredin
The missing speaker labels
A transcription system can return the words of a conversation without telling you which participant said them. Hervé Bredin came to that problem through years of academic research before co-founding pyannoteAI, where he became chief science officer. His open-source pyannote.audio toolkit focuses on speaker diarization: tracking which speaker is active in a recording. When OpenAI released Whisper, users found a natural pairing. Whisper supplied the transcription; pyannote supplied the missing speaker tags. Bredin connects an inflection in the toolkit’s adoption to that release.
The adoption story comes with a birthday request. With his forty-fifth birthday a week away and the project approaching 10,000 GitHub stars, Bredin asks the audience to give it a star as a present.
The underlying distinction is simple. Speech-to-text, or STT, maps recorded or streaming audio to a sequence of words. It answers what was said. But a conversation flattened into one text stream can be hard to interpret: the words alone do not tell you where one participant’s contribution ends and another’s begins.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From what was said to who said it
Speaker-attributed transcription associates the words with speaker tags. That changes the question from what was said to who said what. For a meeting assistant, this may provide enough structure to summarize the discussion and, once the tags are associated with participants, assign an action item to the person who actually accepted it. A correct sentence attached to the wrong participant can produce an incorrect meeting record.
The same need appears in several applications:
- Video translation and dubbing: translated speech needs a consistent voice for each participant. Switching voices between unrelated speakers breaks that correspondence.
- Meeting notes: commitments and contributions need to remain attached to the right people.
- Podcast intelligence: finding the same guest across episodes—or across different podcasts—requires connecting speech to a recurring person.
These applications motivate speaker information, but they ask for different amounts of it. Maintaining consistent labels inside one recording is a narrower task than recognizing a person across recordings.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Timing and delivery change the interpretation
Even correct speaker attribution leaves information out. In Bredin’s example, the timing reveals that the black speaker interrupts the green speaker. A sequence of speaker-labeled sentences without precise timing can conceal that interruption. It can also conceal a backchannel: a short acknowledgment made while someone else continues speaking. Such a response may indicate agreement without taking over the conversational turn. Missing it can remove a small but consequential contribution.
Pauses matter for the same reason. A gap between turns can affect how a response is interpreted, even though it contributes no words to the transcript. The next layer is delivery: how something was said. Bredin offers laughter and coughing as examples of sounds whose possible meaning depends on context. A laugh might respond to a joke or to something awkward; a cough might accompany discomfort. These are possible interpretations, not reliable labels for someone’s state of mind.
Prosody—patterns of stress, rhythm and intonation—and disfluencies add further information. Bredin’s example is “The dog ate the cake”: a child can use the same words while conveying different implications through emphasis. Preserving that information would give a downstream language model, or another analysis tool, more to work with than the word sequence alone.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Who is being addressed, and where?
A broader view asks who is talking to whom. Addressing an entire audience is different from answering one person’s question during Q&A. The acoustic environment adds another layer: the same exchange in a quiet room and on a street comes with different contextual information. These are parts of the conversation-understanding problem Bredin wants to pursue, but he narrows the technical discussion to speaker diarization. Detecting who speaks when does not, by itself, solve addressee detection, prosody or intent.
To illustrate demand for that narrower task, Bredin shows a Hugging Face audio-model ranking sorted by downloads. In the displayed snapshot, three of the top seven models concern speaker identity or diarization; the remaining models concern speech-to-text. It is a historical illustration of how often developers need speaker information alongside words.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
From speech activity to consistent speaker labels
Bredin builds up diarization through three increasingly specific tasks:
- Voice activity detection: identify the regions where anyone is speaking.
- Segmentation: locate speaker changes within those regions and find overlapping turns, including interruptions and backchannels.
- Speaker assignment: give the turns consistent speaker labels throughout the recording.
The fine detail in segmentation matters. A very short affirmative response may carry essential information about a participant’s reaction. Detecting a long region of speech while losing that small turn is not enough. Conversely, finding the turns is not yet full diarization: the system must also establish which turns belong to the same speaker.
The number of speakers is generally unknown in advance. A meeting’s attendee list can provide guidance, but it is not an exact inventory of voices: two people may share a channel, or someone who was not invited may join. Unlike classification against a fixed list of known classes, the system must infer how many distinct speakers are present.
Those classes are also anonymous. Diarization normally returns labels such as speaker one, speaker two and speaker three, not John, Hervé and Jack. The labels are interchangeable as long as their assignment stays consistent. In the two-speaker diagram, changing every green label to black and every black label to green leaves the diarization equivalent. What matters is the grouping of speech by speaker, not the spelling or color of the label.
Unknown speaker counts and interchangeable labels are only part of the difficulty. Speakers overlap, some turns are extremely short, and speaking time can be highly uneven. A participant who contributes only a brief acknowledgment gives the system much less speech to work with than the person leading the discussion. Noise and other acoustic conditions add the difficulties familiar from speech processing more generally.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Inspecting diarization errors in the notebook
The first Python notebook demonstration makes these errors visible. Its input is a roughly thirty-second telephone conversation between two women. After an uncertain exchange of greetings, they introduce their locations—New Jersey and Texas—and discover that both are originally from Chicago. Bredin presents a manually labeled reference for the sample, then applies the open-source Community-1 diarization model.
The local workflow is to download the model from Hugging Face, use PyTorch’s MPS backend on his Mac, and apply the pipeline to the audio file. After switching from an already processed notebook tab to the live run, Bredin compares the resulting prediction with the reference annotation. The aligned timelines make it possible to inspect where the output diverges from the expected speaker activity.
The comparison separates three kinds of error:
| Error | What went wrong |
|---|---|
| Speaker confusion | Speech was assigned to the wrong speaker. |
| False alarm | The system detected speech where the reference contains none. |
| Missed speech | Reference speech was not detected. During overlap, this includes detecting only one of two active speakers. |
That last case is easy to overlook. A system can correctly detect that someone is speaking and still miss part of the conversation because another person is speaking at the same time.
Bredin uses pyannote.metrics to combine these errors into diarization error rate, or DER. The quantities are durations, not word counts:
Bredin reports about 5% DER for Community-1 on this telephone sample.
Editor: The metric’s scoring convention counts speaker-time, not elapsed file time: one second containing two reference speakers contributes two speaker-seconds when overlap is scored. Comparisons also depend on whether overlap is included and on the boundary collar, a tolerance around reference boundaries.
He next submits the same sample to the cloud API for Precision-2. In the displayed result, it fixes a speaker confusion and reduces false-alarm and missed-speech errors. Bredin reports about 3% DER for Precision-2 on the same sample. These are results from the short demonstration recording, not dataset-wide benchmark scores.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Accuracy depends on the recording
How well does diarization work beyond that short call? Bredin’s answer depends on the use case. He reports a best-system DER of about 8% for conversational telephone speech between two people. For a noisy restaurant conversation involving several friends, he reports about 41% DER even for the best system. These are separate scenario-dependent comparisons, not additional scores for the two models in the notebook. The number of participants and the acoustic environment change the difficulty substantially; noisy group conversation remains far from solved.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why a leaderboard score may not transfer to a meeting
Once transcription and diarization are available, producing speaker-attributed text might appear to require only assigning a speaker to each word. The first obstacle is that recognition itself may already be wrong. Bredin argues that most STT models are trained on single-speaker data and generalize poorly to multi-speaker recordings with overlap and speaker changes. His example uses NVIDIA Parakeet and the Open ASR leaderboard, where ASR means automatic speech recognition.
For Parakeet on the AMI meeting dataset, he contrasts the leaderboard’s word error rate, or WER, with his team’s result:
| Evaluation described by Bredin | Recording condition | Reported WER |
|---|---|---|
| Displayed Open ASR leaderboard | Individual headset microphones, giving single-speaker speech | 11.4% |
| His team’s AMI evaluation | A distant microphone in the middle of the meeting table, capturing multiple speakers | 26% |
The dataset name is the same, but the input conditions are not. Bredin explains that the headset recordings and the shared table microphone present materially different recognition problems. The comparison changes both microphone placement and conversational conditions; it does not isolate the effect of overlap alone.
The difficult conditions include distant speech, speaker changes, crosstalk and interruptions. Bredin also raises code switching, changing languages within a sentence. For a meeting application, the practical evaluation rule is to test the kind of recording the product will actually receive. A score from isolated headset speech does not establish performance on a shared room microphone.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Speaker attribution is not a simple timestamp join
The second obstacle is reconciliation: making the transcription and diarization agree about which words belong to which speakers. Their timestamps can disagree. Their coverage can disagree too: diarization may detect speech that STT never transcribes, or STT may produce words where diarization reports no speech. During overlap, two active speakers may be represented by only one recognized word stream. A timestamp join cannot assign words that recognition never produced.
In the second demonstration, Bredin applies Parakeet to the same telephone sample. It returns a sequence of words with corresponding start and end times, providing the word intervals needed for attribution.
Editor: The published demo implementation identifies the variant as
nvidia/parakeet-tdt-0.6b-v3and requests timestamps withtimestamps=True.
He places those words beside Precision-2’s speaker turns and examines the assignments one by one:
| Displayed case | Why attribution is easy or difficult |
|---|---|
| A word interval is shifted slightly left | The yellow speaker still appears to be the clear match despite the timing mismatch. |
The third word, oh, falls between two diarized turns | The word has no unambiguous turn to attach to. Bredin replays it while considering the assignment. |
The word okay coincides with two active speakers | Diarization offers two candidates, but transcription supplies only one word. |
These are different failure modes: imperfect boundaries, a gap between turns, and competing speakers during overlap. Treating all of them as the same interval-matching problem hides the uncertainty that reconciliation must resolve.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Preserving the overlapping exchange
Bredin then submits a job to pyannoteAI’s cloud service that combines Precision-2 diarization, Parakeet transcription and reconciliation. The returned output contains words assigned to speakers, so the application does not have to resolve each conflicting interval itself.
The revealing example is the transition between the two introductions. One participant’s trailing um overlaps the other participant’s opening words. Bredin plays that passage and points out that the output interleaves words from the two speakers, retaining the filler as the other person begins speaking. He reports that the system gets this overlap right. The demonstrated success is specific: it preserves both contributions in this exchange, rather than losing one inside a single speaker’s turn. The demonstration does not reveal the complete reconciliation algorithm.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The tools behind the demonstration
Bredin closes the demonstration by pointing to the talk tutorial in pyannoteAI/tutorials. Its components have distinct roles: pyannote.audio supplies open-source diarization, pyannote.metrics evaluates the predictions, and iPyannote provides the interactive visualization widget. The initial i, he explains, stands for interactive. The SDK provides access to the premium models, complementing the local open-source workflow shown earlier.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Exclusive diarization and independent STT models
An audience member asks where conflict resolution happens: is it learned during model training, or supplied by heuristics around the models? Bredin says the complete reconciliation technique is proprietary, but identifies one disclosed part available in Community-1: exclusive diarization.
During overlapping speech, Bredin explains, exclusive diarization selects the speaker considered most likely to be transcribed by the STT model. This removes a competing speaker assignment from that view and makes reconciliation simpler. It is one major part of the approach, not a complete explanation of how the earlier demonstration recovered and interleaved the overlapping words.
Editor: The Community-1 model card documents separate
output.speaker_diarizationandoutput.exclusive_speaker_diarizationresults. The exclusive output is a simplified view for reconciliation; it does not mean the original recording contained no overlap.
Pressed on whether this requires changes to model training, Bredin emphasizes the intended boundary: the orchestration should work without modifying the STT model itself. He describes support for different transcription systems—including privately fine-tuned models—as a design goal. That lets a team retain a recognizer adapted to its own use case while adding diarization and reconciliation around it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The exact tutorial for the two demonstrations, including diarization error inspection and speaker-attributed transcription. Full reproduction requires model access and cloud credentials.
Model card for the open-source local pipeline, including its exclusive-diarization output for reconciliation with transcription.
Reference for DER components, speaker-label matching and evaluation settings used to interpret the first demonstration.
Official model card for the exact Parakeet variant found in the talk's tutorial implementation.
Open-source toolkit for voice activity, speaker segmentation and speaker diarization.
Read the complete timestamped transcript
- 0:00
[upbeat music] Uh, good morning, everyone.
- 0:16
Thanks for, uh, being here to the, uh, Voice and, uh, Vision session. Uh, so I'm, uh, Hervé Bredin, chief science officer and co-founder at pyannoteAI. Uh, so I'm going to talk to you, uh, today about, um, conversations, understanding conversations and, uh, uh, what you can do on top of, uh, transcription.
- 0:37
Uh, so a quick words about myself. So I've been a, an academic researcher all my life until, uh, two years ago when I, uh, started this company. Uh, oops, sorry.
- 0:48
Uh, what happened? Yeah. Uh, so basically, yeah, I, I, I worked on this topic called speaker diarization, which I'll, uh, introduce a bit later for those of you who don't know this, uh, weird, uh, word that is tricky for me to pronounce.
- 1:03
Uh, but basically, over the years, I built an open source toolkit called, uh, pyannote, uh, which focuses on speaker diarization, and that became quite popular over the years, in particular since, uh, uh, OpenAI, uh, uh, released, uh, Whisper speech-to-text, uh, open source models.
- 1:19
Basically, Whisper was, uh, some kind of revolution, uh, in terms of STT. Uh, the fact that it was free, that it was, uh, very good and, uh, but it didn't provided the actual, uh...
- 1:28
didn't provide the actual names and tags of the speaker, so people naturally turn to pyannote to combine the two. And you, you can see that on the, uh, inflection points on the GitHub star history.
- 1:39
And, uh, just a few words, I happen to turn [REDACTED:age] uh, uh, in just one week from now, and we are almost at 10K, uh, stars on the GitHub.
- 1:48
So please, uh, give me a, a birthday gift by, uh, just, uh, going to the website and, uh, to the GitHub website and add your star. That will make my, uh, birthday.
- 1:57
Um, so, um, let, let, let's go to the, to the core of the presentation. So you all know what transcription or a speech-to-text is. Basically, that's answering the question, what was said from the recording or streaming of a, a conversation.
- 2:12
Basically, you get from the audio, and you get a sequence of words, uh, uh, as output. Uh, that's, for instance, what Whisper does. Uh, but, uh, without the actual, uh, names of the speaker, the usually the conversations on the transcript are not really, uh, understandable.
- 2:29
So the next step after transcription is basically to attribute, uh, a, a speaker tag to each, uh, of the words. So that's, uh, I call here a speaker attributed transcription.
- 2:39
And so that, that answers the question, who said what? But for some application, that's enough. Like for instance, uh, uh, for maybe for meeting note takers to make a summary of a, of, of a meeting, that might be enough to assign an action point to the right people to, to know exactly that, uh, Hervé said that and,
- 2:57
uh, John said, said th- this other thing. But we-- to understand the conversation, uh, a bit better, uh, we, we might want, uh, uh, to go, uh, slightly, uh, better.
- 3:09
Uh, sorry, I forgot about this slide. So basically, uh, there are many cases where, um, knowing who said what is as important as what was said, actually. So I have here a few examples.
- 3:19
Like, for instance, uh, video dubbing, automatic video dubbing. Um, uh, s- knowing who said what is actually really important to put the actual, uh, uh, good correct voice to the, uh, right speaker.
- 3:33
So i- if you want to automatically translate, uh, um, a video from one language to another, you want the voices to be consistent as well. So for that, you need, uh, to know who's, uh, who, who spoke when.
- 3:46
Same for meeting note takers, and there are other applications like, uh, for instance, podcast, uh, what I call podcast intelligence, like being able to track a speaker across different, uh, uh, episodes of a podcast or, uh, a- across different podcasts, finding, uh, the same guest across, uh, multiple podcasts.
- 4:05
But yeah, so, uh, I was a bit, uh, o- over, uh, just before. But, uh, the, just knowing who said what is sometimes not enough to really understand the conversation, the conversation, how it goes.
- 4:16
So knowing who said what and when actually, uh, brings even more information. For instance, in this example, um, uh, without knowing exactly when each words were pronounced, you, you can't detect that, uh, actually, the black speaker interrupted, uh, the green one.
- 4:32
You can't really detect that, uh, maybe this small backchannel that, uh, the, the, the green, the, sorry, the black speaker, um, uh, does during the speech of the, uh, of the green speaker.
- 4:43
If you miss it, basically, you, you don't really understand that, uh, maybe the black speaker is actually agreeing, nodding, uh, at what the, the green speaker says. So knowing exactly when this small word, uh, has been pronounced, uh, is very important.
- 4:58
And same, without precise timestamps, you can't really know whether I made a... I make a, a pause in between two, uh, speech terms, which might, uh, convey additional information about, uh, my state of mind, my, I don't know, the, the, the point that I'm trying to, to make in the conversation.
- 5:14
And we can go even a step further. Like we, we might want to know who said what and when and how they said it. Like, uh, I have a few examples here.
- 5:24
If, uh, some of you are laughing at some point, is it because, I don't know, I said something stupid or I was actually funny? Uh, same for coughing. Uh, maybe there's something in the conversation that makes me, uh, uncomfortable and, uh, I'm coughing for some reason to, uh, to hide be- behind that.
- 5:41
And, um, yeah, so I also mentioned stress, disfluency, prosody. Um, the fact that, uh, I, uh, stress on particular words in a sentence might actually change completely the meaning of the sentence.
- 5:54
Uh, I have this example in mind, like, uh, I don't know, in the sentence, "The dog, uh, ate, uh, the cake." If, uh, I don't know, you're talking to a child and, uh, the, the child tells you, "The dog ate the cake," or, "The dog ate the cake," or, "The dog ate the cake," uh, the, the meaning
- 6:11
might be slightly different. So, uh, stressing particular words, uh, being able to have a, a voice AI actually understand the stress and, uh, and, um, this kind of, of, um, low-level details might bring some more, uh, information to the,
- 6:26
to any downstream LLMs or, or whatever tools that you use after this first, uh, enriched transcription step.
- 6:35
And if we take even a step, uh, a step back, uh, um, knowing who's, uh, talking to who. Uh, am I, uh, here in this example, I'm probably can be seen as this, uh, gray speaker addressing to all of you.
- 6:49
But at some point, if, uh, uh, at the end of this talk, I, uh, you do have questions, then, uh, I will be probably the, the green one talking to one particular person answering the particular question.
- 6:58
So knowing that, having a higher level view of the conversation, it also opens up a, a, a, a new bunch of, um, applications.
- 7:09
And all of this, uh, also happens in a acoustic environment that gives you information about the context, whether it's in this quiet room, uh, for now, uh, or if I'm in the street.
- 7:22
Uh, this bring more context, uh, to, to, to the game, and you can actually, uh, take a, a better decision about, uh, what's going on, why a person is actually saying what, what they are saying and, and et cetera.
- 7:34
So all of this is basically the, what we are trying to work on at, uh, pyannoteAI, but we, uh, started with a, a smaller problem called speaker diarization, and, uh, that's really what I want to talk a-about, uh, today.
- 7:50
So speaker diarization is the... I, I... Yeah, i-is, is usually... Yeah, what I, what I wanted to say about this slide is that, uh, who said what is, uh, usually just as important as what was said, and you can tell by instance by going through, uh, to the Hugging Face, uh, uh, model repository and filter, um, all
- 8:09
the models there, uh, with the ones that have the audio tag. Uh, where is my mouse? Yeah, here. So this is what I did, and then you sort them by, uh, by number of downloads.
- 8:19
And basically, uh, the... among the, the top, uh, here, seven models that, uh, are, are at the top of this list, three of them, the, the, the first three are actually related to speaker identity and speaker diarization, and the, the, the other ones are STT.
- 8:33
So that, that's one way of showing that, uh, speaker diarization is actually quite, uh, quite important on, on, i-in the community.
- 8:42
So, uh, going, uh, into a bit more detail about what speaker diarization is. So, as I said, it's answering the question who speaks when. So starting from the recording of a conversation, uh, basically, the first step that people usually do is, uh, basic voice activity detection.
- 8:59
So you basically tell whether someone is speaking at, at one time or not. Anyone is speaking at, at one time or not.
- 9:06
And then you can go a step further. Uh, you can actually segment those speech regions into smaller speech turns. Um, so thi-this is what I call segmentation, and it, it includes finding, uh, speaker change points between, uh, i-into those long, uh, in those long speech regions, you might find a, you know, a speaker change point, and, uh,
- 9:26
also find regions where, uh, people are actually interrupting each other. Like, uh, in, uh, in the first example here, there's definitely here someone interrupting, uh, someone else, and it's possibly here some kind of back channel where someone is actually, uh, uh, nodding or, uh, saying, "Mm-hmm, okay," th-this kind of stuff.
- 9:45
And this is the kind of small speech turn that, uh, you don't want to miss because sometimes they actually convey, uh, the most important information of the conversation. If it's a small yes here, uh, and if you miss it, uh, then, uh, you, you have no idea what's the, the, the state of mind of this speaker.
- 10:05
But that's not yet speaker diarization. Speaker diarization goes all the way, uh, to actually assigning a speaker identity to, uh, each, uh, speech turn.
- 10:16
And, um, so here in this example, the, the speaker diarization system automatically detected that there are two speakers, the green one and the, uh, black one. Um, but, uh, we usually don't have any prior knowledge on the actual number of speakers in the conversation.
- 10:33
You might have some kind of, uh, uh, guidance on, uh, I don't know, if, if you are a meeting note-taker, for instance, uh, you, you have the list of attendees.
- 10:41
Uh, so you might have an idea of the number of people that were invited, but that doesn't mean that, uh, someone, um, two people join from the same channel or, or, uh, that an attendee that was not invited actually, uh, join in finally and, uh...
- 10:56
So, uh, that, that makes the problem, uh, kind of difficult in the sense that we don't know in advance the number of classes that we're supposed to detect, uh, contrary to other, uh, let's say, classical machine learning problems.
- 11:08
And also, we don't know the identity of the speaker, uh, that we're supposed to, uh, uh, diarize, in the sense that speaker diarization does not really output, uh, John, Hervé and, uh, and, uh, and Jack.
- 11:23
Uh, but basically, it's more like speaker one, speaker two, speaker three, and they can be permutated, and it's still correct. For example, i-in this example, if I just change, uh, green into black and black into green, the diarization is still exactly the same, uh, from the point of view of the evaluation metrics that I'm going to talk
- 11:40
about and, uh, uh, from the point of view of the actual, uh, speaker diarization task.
- 11:46
So that makes the, uh, this problem, uh, kind of, uh, difficult, and that's why even though, uh, the community, uh, has been working on this topic for a long time now, it's still not solved.
- 11:56
Uh, there are other reasons, like the fact that we have to detect overlapping speech, uh, to handle overlapping speech, uh, that we have to, uh, uh, take into account very short speech turns, that don't have to take into account- Imbalance between, uh, the speech time of multiple speakers in the, in the conversation, all of that makes the
- 12:14
problem difficult. On top of the usual, uh, acoustic conditions, uh, problems and everything that any speech processing, uh, um, um, task, uh, have to deal with.
- 12:27
So, uh, I'm going to switch to a demo to, to show you, uh, a bit more how the, uh, diarization system I evaluated. Um, maybe when you, you, you stumble in- into benchmark, you, you find this, uh, DER acronym, which is, uh, diariz- stands for diarization error rate.
- 12:42
And, uh, so hopefully, um, this will work. So, uh, I have here, uh, a Python notebook, uh, that, uh, I've prepared for, for this talk. It's actually already, uh, available on GitHub.
- 12:56
I'll show you the, the QR code at the end of the presentation, so you can play with it. But so the... in this example, I have a conversation between two [REDACTED:gender], uh, talking over the phone.
- 13:05
Hello?
- 13:06
Hello.
- 13:07
Oh, hello. I didn't know you were there. [laughs]
- 13:08
Neither did I.
- 13:09
Okay. Wait, I thought, you know, I heard a beep. This is Diane in New Jersey. Um-
- 13:13
And I'm Sheila in Texas, originally from Chicago. [laughs]
- 13:16
Oh, I'm originally from Chicago also. I'm in New Jersey now, though. [laughs] Okay. So you, you get the idea, a 30-seconds conversation between two [REDACTED:gender]. So this is the expected output of a perfect speaker-attributed, uh, transcription.
- 13:27
This was manually labeled. And now, uh, what I'm going to, to, to do this here is, is actually run a piano open source, uh, Community-1 model on, on the same file.
- 13:37
So it's actually, uh, running, uh, right now on my Mac. So that's why here I'm using this, uh, MPS, uh, PyTorch, uh, backend. So how it works is that, uh, you first download the model from, uh, Hugging Face with this, uh, line of code, then simply, uh, uh, apply it on, on the audio file.
- 13:54
So now, uh, it's finished and the... there's the, this prediction variable that contains, um,
- 14:01
uh... Sorry, I should have done that on the other tab because this was the one that was already processed. I wanted to do it live, so let's do it live now.
- 14:10
So it's running. And so the, the next step is to actually visualize the errors that, uh, this system, uh, made. So at, at the top here, you have the reference annotation, so the expected output of diarization.
- 14:22
And here is the output of, uh, Community-1, uh, speaker diarization, uh, pipeline, uh, that is open source and free to use. So it makes three kinds of mistake, as I was saying.
- 14:32
Uh, like a confusion. Here, it got the speaker wrong. Uh, it can also make, uh, false alarms. Uh, so false alarm is when it detects speech when there's actually no speech in the ground truth.
- 14:43
And, uh, misdetection is the other way around. And misdetection can actually, uh, happen, for instance, here during overlapping speech, when we actually detect only one of the two speakers here.
- 14:55
And then once we have these, uh, uh, basically errors of, uh, false alarm, confusion and, and misdetection, we, we can actually compute the diarization error rate. Uh, so this is with this library called pyannote metrics.
- 15:08
And, uh, basically, it gives us, in this example, a five percent diarization error rate, which is basically the sum of the confusion, uh, uh, false alarm and misdetection divided by the, the total, uh, duration of, uh, of speech in, in this, uh, in this file.
- 15:25
And, um, so w- we have this other model that I'm currently running. Hopefully it work, it, it works. So it's running, uh, right now on our, uh, cloud API.
- 15:36
Uh, basically, it's, it's, uh, uh, a better model, that Community-1, that we call Precision-2, that is, uh, uh, running here. Hopefully it will work, but as for every demo, it might fail.
- 15:48
If... No, it worked. And, uh, this is the kind of output that you can get. So you see that, uh, it got the, the s- the speaker right here and, uh, makes, uh, slightly less mistakes in terms of mi- uh, false alarm and misdetection.
- 16:00
And overall, in this example, I think we got a, a three percent, uh, diarization error rate here.
- 16:07
So let me switch back to the, to the presentation. Uh, so this was about, uh, benchmarking, uh, speaker diarization model, and I'm often asked how well the state-of-the-art speaker diarization works today.
- 16:19
And it- it's a difficult question to answer because really, depending on the use case, it might, uh, vary completely. For instance, in this example, we have a here conversation telephone speech like the one we just, uh, listened to, like two person over the phone, and we, we can go down to eight percent, uh, diarization error rate.
- 16:35
The best system, uh, does that. But if, if you are now in a restaurant with many friends, with lots of background noise, uh, even the best system reaches like forty-one percent diarization error rate.
- 16:46
So it's far from being a solved problem, uh, but, uh, we, we are working on that.
- 16:52
And, um, so once you have speaker diarization and you have transcription, shouldn't be that hard to actually find a speaker-attributed transcription, right? It's just a matter of, uh, you know, assigning a word to a speaker, a speaker to a word.
- 17:06
But actually, that's not that, uh, easy. And I'll, um... There, there are many reason why that's not that easy. The first reason is that most, uh, speech-to-text models are actually trained on single speaker data.
- 17:19
As soon as you apply them on multi-speaker data with overlap, with speaker change, with all kind of mess, they completely, uh, uh, they fail miserably. So for instance, when you look at the open ASR leaderboard from, uh, Hugging Face, if we look, for instance, at NVIDIA Parakeet that we are using, uh, that, that I will be using
- 17:36
in a, in a demo right after that, they report eleven point four, uh, percent word error rate.
- 17:43
When we apply the very same model on the same AMI, um, on, um, on our side, we, we get twenty-six percent. So the question is, w- why is there a difference?
- 17:55
Actually, it's in the way the, those benchmark are set up here on the, uh, uh... For, for example, in this AMI dataset, that's, uh, uh, a meet, um, a datasets with meetings, uh, between four to five people, I think, in meeting rooms.
- 18:09
And they are basically microphones like the one I'm, I'm, uh, having here, so a headset microphone, as well as one microphone in the middle of the, the table. And, and, uh, those numbers here, uh, on the OpenAI ASR leaderboard are based on the headset microphone.
- 18:24
Those numbers here, uh, are based on the microphone that is in the middle. And so on one side, you have a single speaker speech. On the other side, you have multi-speaker and even distant microphone speech.
- 18:34
So that's why it degrades, uh, a lot. And so that, that, that's... that was my point of the, of this slide. So, uh, when doing speaker attributed transcription, the, um,
- 18:44
the reason why it, it might go wrong is either because STT doesn't work great and, um, uh, it's usually the case that, uh, they don't generalize, uh, very well to multi-speaker recordings for, as I was saying, distant microphone, uh, speaker change, crosstalk, interruptions, you name it.
- 19:02
Uh, maybe code switching as well when you change language in the middle of a sentence.
- 19:07
But also it's, it might be because of the actual reconciliation between diarization and, uh, STT timestamps. Though it may sound, uh, uh, obvious how to do that, as I was saying, I will hopefully show in this, uh, live demo once again that, uh, it's, it's not such a, an easy problem because, uh, STT does not transcribe overlapping
- 19:28
speech well, uh, because, uh, the timestamps disagree between STT and, uh, diarization, and because sometimes diarization will detect speech that the transcription will not transcribe and the other way around.
- 19:43
So, um, second part of the demo, uh, about speaker attributed transcription. So, uh, first, I'm going to, uh, apply, uh, this Parakit model, uh, from NVIDIA that does, uh, by the way, a great job at, uh, transcribing, uh,
- 19:58
uh, uh, th-this sample. And so, uh, Parakit gives us this kind of output. So we have this sequence of words, and for each word, we actually have, uh, the, the corresponding timestamps.
- 20:11
Okay? So, uh, I, I can play it quickly. Hello?
- 20:15
Oh, hello. I didn't know you were there. [laughs]
- 20:16
Neither did I.
- 20:17
Okay. I thought, you know, I heard a beep. This is Diane in New Jersey, um-
- 20:21
Well, you, you get the idea. And then, um, the question is, um,
- 20:28
now we have our best diarization. This is Precision-2 output here. This is Parakit output here. And the question is, uh, now we need to assign a, a speaker to each word.
- 20:37
So le-le-let's do it, uh, step by step, for instance. Uh, so in this example, even though the timestamps of the, the words, uh, are a, a bit shifted, uh, on the left here, it's obvious that, uh, it's probably the yellow speaker who spoke the, the gray, uh, word.
- 20:53
But if you go just the third word, so now we have this. So, uh, we have this, uh, word, uh, the word O here, uh, that is, uh, in between those two, uh, speech turns according to diarization.
- 21:05
Which one do you assign it to? Um, so we, we... I, I can listen to it. Oh, hello.
- 21:11
Oh, hello. Right. It's not, it's not quite, quite sure. And there are many other, uh, positions where it happens here same. There's a word here, okay, that, uh, apparently according to diarization, there are two speakers speaking and, uh, only, only one word from the, from the transcription.
- 21:26
So that, that's why, uh, at, uh, pyannoteAI, we came up with this, uh, new, uh, STT orchestration thing that basically does, uh, its magic and, and does the job for you.
- 21:35
So here I j- I just submitted, um, to our cloud, uh, API, um,
- 21:42
uh, a job to transcribe and, and, and use both Precision-2 diarization model and Parakit transcription model and take care of the... what I call this, uh, reconciliation between the two.
- 21:54
And you end up with this kind of, uh, with this kind of output. Uh, what's nice is that so now you... it did the, the, the job for you, but in particular, I, I like to go to this particular example.
- 22:06
Even where there was overlapping speech before, it actually managed to, uh, interleave the, the words from those speaker. Let me play th- just this part and focus on this, uh, area maybe.
- 22:16
In New Jersey, um-
- 22:17
And I'm Sheila in tech- And there, there's actually a and, uh, that, uh, or this is the um, here that, that is, uh, actually, uh, uh, the, the, the, the, the, the yellow [REDACTED:gender] a- actually, uh, has this um, at the end of their, her speech and, uh, the, the blue one actually interrupts her.
- 22:37
But they, they, they're actually overlapping, and we managed to get them right.
- 22:42
Um, yeah, so th-this is, uh, the end of my talk. I, I, I wanted to, uh, give you a f- a, a few minutes. I have two minutes left for questions, but I wa- just wanted to say that this demo, uh, that I, I did is already available on, uh, in the pyannoteAI/tutorials, uh, GitHub repo.
- 22:58
So in there, you, you, you'll be able to play with pyannote.audio, open source toolkit for diarization, with pyannote-metrics for evaluation, uh, with iPyannote for this nice, uh, visualization, uh, widget that we've, uh, implemented.
- 23:10
I stands for, uh, not iOS, iMac or whatever, but f- really for interactive and, uh, also the SDK to, to play with our premium models. And, uh, I'll stop here, and, um, happy to take, uh, a few questions.
- 23:23
We have one more minute left. [clapping] Yes, please.
- 23:34
Um, so what's the trick, uh, that you use to resolve the conflicts?
- 23:40
So-
- 23:40
Is it part of the model training, or is it some sort of like heuristics that you have around to identify who to attribute the word to?
- 23:48
Yeah. So the question is, uh, what's the trick to, uh, actually solve this problem of a reconciliation? Uh, so that, that's a, a, a proprietary trick. But what I can, what I can say is that we, uh...
- 23:59
there's already part of this trick that is available in the Community-1 model, which we call exclusive diarization. Basically, uh, what we do is that we find a way when there is overlap to actually select, uh, the, the most, uh, likely of the two speaker that will be transcribed by the STT model.
- 24:14
So that simplifies actually the reconciliation between the two. So that, that's one major part of the, uh, of, of the approach.
- 24:21
Part of the model training is like a se- couple of heuristics you have to-
- 24:25
Yeah. So, so the question is about, uh... I, I'm repeating because I've been told I need to repeat questions. Uh, so it's not part of the model training in the sense that we really plan to support any kind of STT without having to change the STT model itself.
- 24:38
So really, it's supposed to work with any STT, uh, even fine-tuned ones that you might have internally for... because you, you fine-tuned them for your particular use case and nobody has it.
- 24:48
Uh, you can combine it, uh, like that.
- 24:50
Yes.
- 24:52
All right. Yeah, eighteen seconds, so I guess we need to stop. I'm sorry. Maybe we can talk offline. Otherwise, uh, they, they'll, they'll beat me, I guess. [laughs]
- 25:00
Thank you very much. I'm sorry. [clapping] [outro music]