Multi-Speaker Diarization
Multi-Speaker Diarization involves the process of segmenting audio recordings by speaker
labels. The goal of this project is to identify human speech and annotate who spoke and when.
You will identify speakers as Speaker 1, Speaker 2, and so on, as the true identity of the
speakers will be unknown.
The intent of the project is the creation of speech-to-speech datasets on human-to-human
conversations (namely, natural dialogues and interactions). For this reason, it is important to
capture and focus on the speakers that have such natural dialogues and interactions, while
ignoring speech that is not representative of conversational data (such as commercials,
advertisements, show-opening clips).
In this project you will perform the following actions:
1. Listen to the audio and identify segments where human sound is audible.
2. Use your best judgement to determine if the human sound is part of a conversation with
another human or not.
If yes, then identify each unique voice and annotate it as a Speaker. If you cannot identify to
which speaker the voice belongs, annotate as Unknown.
If not, then annotate as Ignore.
3. Indicate the total number of unique speakers in the audio.
4. Indicate if the audio file starts in the middle of human speech or sound, and where is it cut off.
Discard and Skip
Discard the task in the following scenarios if they apply to the entire audio clip:
If no speech is present in the entire clip.
If the audio is so badly distorted that it is impossible to determine whether speech is present.
If there are audios that do not load, those tasks will not be presented to you.
If by any chance audios in the platform do not load, DO NOT use the discard option for this
scenario. Try to refresh the browser, check your connection, or contact the Centific team for
more information.
You may skip the task when the audio does not load (rather than discarding).
End-to-End Process
To process a task in this project, follow the steps below:
Step 1: Listen to the audio
To listen to the audio recording, use the audio segmentation tool. Pay attention to the distinctive
speakers you can hear.
There will be no audio reference for speakers, so always refer to the audio tool to listen to the
voices as much as needed to determine if a voice speaks multiple times throughout the audio.
Step 2: Label segments
Use the audio segmentation tool to label segments. Each segment should contain the speech of
a single speaker turn.
If there is a dialogue between two or more speakers, each speaker turn would be in its own
segment. This is also true if the voices overlap: the segments will also overlap to reflect that.
To label the speaker segments, follow the instructions below:
1. Timestamp Annotation: Identify the start and end times of each speaker's tum, including
instances of overlapping speech.
To start the segment, click on the start of where the speaker audio starts and drag the cursor to
the end of the segment, being careful of following all the segment requirements in the
guidelines.
End the segment every time there is a pause of over 2 seconds between the end of a speaker's
last utterance and the beginning of their next utterance.
Carefully note when speakers overlap, indicating the start and end of the overlap by creating a
segment. You must pay attention to the point at which one speaker stops and another begins by
following the described process above.
Speaker Labeling:
Assign a speaker label (for example, Speaker 1, Speaker 2, Speaker 3...) to each segment of
speech. Each speaker turn must still be segmented separately with its own label.
Unknown
Ignore
To perform the Speaker Labeling step, follow the guidance below:
Assigning a Speaker label
. Assign a speaker label (for example, Speaker 1, Speaker 2, Speaker 3...) to each segment of
speech.
Assign a speaker label to the primary speaker or speakers (for example, the primary speakers
of a conversation, the host or hosts of a podcast, the person being interviewed or a guest in a
podcast or TV show episode).
If the voice of the current speaker sounds distinctly different from the previous speaker, assign a
new speaker label.
Once you've assigned a specific speaker a label, that label should always be used for that
voice.
If at all possible, provide a speaker label. There may be cases where you are unsure if a
segment is from a previous speaker or a new speaker. In these cases, go back and listen
carefully to the previous speaker(s) to help determine the label for the current speaker you are
segmenting. Additionally, use the context of the conversation, such as the names used and
voice characteristics, to inform your decision.
Labeling as "Unknown"
In extreme cases of uncertainty, where you are absolutely unable to determine if it is a previous
or new speaker, select Unknown. This might occur due to noise or other audio events. Use the
"Select reason for 'Uknown' segment" dropdown to add details about what
speaker(s) in the main production and may have other differentiating characteristics like volume
or audio quality.
In cases where context supports that the host of a podcast, for example, is doing an ad or
sponsorship, do not select Ignore. Instead, give this segment a Speaker label. This includes
cases where it is a monologue of an ad or sponsorship from the podcast host. These cases may
be ambiguous and require you to use your best judgement to determine if it is the podcast host,
or an outside speaker
In these cases, you can identify the criteria of the audio that caused this label in the Notes field.
If there is an audio sequence that contains multiple speakers or voices, but it is all a part of one
sequence such as background music vocals, intro segments. commercials, ads, and so on,
create a single segment and mark as Ignore.
Similarly to background music and songs, ignore any pre-recorded segments (for example,
[sports event (could contain player voices, crowd cheering) + commentators (1+)]), that do not fit
into the main conversation. Such segments may be in the foreground or background.
Even if it is a pre-recorded segment, you should still segment it, like Example 1 in the Examples
section. The pre-recorded segment is a part of the main line of exchange, even if it is not the
main hosts or narrators speaking (they could refer to the pre-recorded segment and the
speakers in the segment).
In all these cases, you must specify in the "Select reason for 'Ignore' segment" dropdown what
type of element is being marked as Other
Background music with vocals
Intro song/audio
Intro segment with various speech segments
In summary: tag all speech that is part of the main narrative or conversation, regardless of
whether the voice is natural, robotic, or fictional, and only Ignore isolated, unrelated commercial
content.
If you select Unknown or Ignore, you should add Notes in the appropriate text field in the Ul. You
are encouraged to use them for the following:
For Unknown labels, input any relevant elements or details that impacted your ability to identify
a speaker
For Ignore labels, identify the criteria which were met that resulted in this label.
Handling pauses
A segment can overlap completely with another speech segment (for example, speaker 1 is
talking, speaker 2 backchannels; speaker 1 does not make a pause longer than 2 seconds).
You must use careful judgement to identify if the pause is significant or not. A pause is
considered significant if the speaker pauses for over 2 seconds. If they simply speak slowly,
make a pause due to an interruption for under 2 seconds, or to think about their next word, then
the pause for under 2 seconds, none of this significant and should not be considered as a new
segment.
For further details on the Speaker Labeling steps, see the Best Practices section.
Handling "fantasy" or gibberish language
In some cases, you may encounter speech that does not represent a real-world language but
still appears in the audio. This includes fictional languages, gibberish, or robotic sounds that are
not "speech" in the human sense. Handle these cases as follows:
Fictional or fantasy languages (e.g., Klingon, Elvish, Dothraki):
If you can determine the speaker, label the segment normally with a Speaker label.
Do not mark these as foreign languages.
In the UI, select "No" for "Was a foreign language present in the audio?". You do not need to
select any language from the list.
Gibberish speech (e.g., fake-Italian nonsense words, playful "babble" sounds):
Treat these sounds as backchannels or feedback-like vocalizations.
If you can determine the speaker, label the segment normally with a Speaker label.
D Do not mark these as foreign languages.
In the UI, select "No" for "Was a foreign language present in the audio?".
Non-linguistic robotic sounds (e.g., R2-D2 beeps, tones, or other purely electronic signals):
Examples:
Example 1: An English interview is dubbed into Spanish. You hear the English speaker faintly in
the background, but a Spanish voice has been added on top.
What to do: Segment only the Spanish dubbed speech. Ignore the English underneath. Select
"Yes" for the option "Did the audio contain dubbing?".
Example 2: A French news clip is dubbed into French, with the English speaker's voice still
audible behind the English translator.
What to do: Create segments for the French dubbed speech, assign speaker labels as usual,
and select "Yes" for the option "Did the audio contain dubbing?".
Dubbed audio where the working language is underneath
In some rare cases, the working language (the language you are assigned to annotate) may be
the original base language, while a foreign dubbed track is placed on top of it.
In these cases, focus on labeling only the working language, even if it is faint or partially
covered by the dubbed audio.
Do not segment or label the foreign dubbed layer.
Apply speaker labels as you normally would for the working language.
If the working language is difficult to hear, do your best to capture and label it accurately.
Example:
If you are working on Spanish and the original Spanish speech is faintly audible under an
English dub, segment and label the Spanish voices only. Do not label the English dub.
Reminder: Always prioritize labeling the working language. The foreign dubbed layer should be
ignored.
NOTE-Don't mix up "foreign language" vs. "dubbed audio"!
These two cases look similar at first, but they're not the same thing:
Foreign language audio: you've got people naturally speaking more than one language in the
file (e.g., the task is in French, but someone suddenly speaks in Spanish). In this case, you
segment and label both (Spanish, German, etc.), mark "Yes" for "Was a foreign language
present in the audio?", and pick the correct option from the dropdown (or "Other" /
"Indeterminable" if needed).
Dubbed audio: the target language is layered on top of a different language (e.g., a Spanish dub
over faint English in the background). In this case, you skip labeling the background foreign
language and only segment the dubbed voice in your language. Then you select "Yes" for Did
the audio contain dubbing?
The fundamental difference:
If someone is actually switching languages in the conversation, it's foreign language audio.
If a voice in the target language has been recorded over another language, it's dubbed audio.
Keep this in mind so you don't over-label or miss the dubbing flag!
Step 3: Indicate the total number of unique speakers
After you have completed the audio segmentation, indicate the total number of distinct speakers
labeled as "Speaker 1," "Speaker 2," "Speaker 3," etc., across the entire audio.
Only include segments that have been assigned a specific Speaker label. Do not count
segments labeled as "Unknown" or "Ignore" in your total.
Enter the total number of distinct speakers in the text box How many unique speakers did you
identify in the whole audio? Only enter whole numbers in the text box.
You will always select two options in this step:
If the audio begins normally but cuts off at the end Start: No, End: Yes
If the audio is cut off at the beginning but ends normally Start: Yes, End: No
If the audio is cut off at both ends - Start: Yes, End: Yes
If the audio is complete at both ends Start: No, End: No
Step 5: Review and submit the task
The final task output should include the start and end timestamps for each speaker's turn
(including overlaps) and the associated speaker labels, indications of whether audio cut off and
where, as well as any notes for the Unknown/Ignore labels.
Best Practices
To annotate a task in this project, use the following best practices:
1. Speaker labeling
1.1 Mark the speaker index for each segment. The first speaker of the segment will be Speaker
1, the second unique speaker will be Speaker 2, and so on.
1.2 Do not cannot overlap segments of the same speaker with each other as, logically, a person
cannot talk at the same time as themselves (i.e., you cannot overlap a segment marked
Speaker 2 with another segment labeled Speaker 2). If you overlap segments of the same
speaker, either check the label (if they are different speakers) or merge the segments (if they
belong to the same speaker).
2. Language handling
If you hear a different language in part of the audio, label those segments as usual. Use
"Unknown" if you can't identify the speaker. Select "Yes" for "Was a foreign language present in
the audio?" if any foreign language appears.
IMPORTANT! If the whole audio is in the wrong language, discard the task without labeling.
3. Overlapping speech
3.1 Pay close attention to instances of overlapping speech, where two or more speakers are
talking simultaneously. In these cases, note when one speaker stops talking and another starts,
even if the overlap occurs.
When you hear overlapped speech between multiple speakers (for example, Speaker 1 and
Speaker 2), note down when Speaker 1 stops speaking and Speaker 2 starts speaking.
Speaker 2 can start speaking earlier than when Speaker 1 can stop speaking.
4. Accuracy of labels and timing
4.1 Ensure that start and end times, as well as speaker labels, are accurately and
consistently applied. Such as, Speaker 1's voice should always receive the Speaker 1 label.
4.2 Provide clear and comprehensive annotations that accurately capture the conversational
dynamics, including all speaker turns and overlaps.
4.3 Speaker utterance boundaries (start and end time) should be accurate to 250 ms precision.
5. Labeling non-verbal speech
5.1 Label the following verbal sounds as long as they are a part of a natural dialogue or
interaction between speakers:
Laughter from another person, as this is feedback for what another speaker is saying
Backchannel is when the person listening produces audible vocal activity to signal active
listening, such as: uh-huh, mmmm, and all other sounds that show active listening by providing
relevant feedback to what is being said, without using words (they can show approval,
disagreement, encouragement, and so on).
5.2 Label audible and obvious breathing with the corresponding speech when applicable.
The main focus of segmenting is to capture conversation. If there is no audible and obvious
breathing after speech, do not search for it. If there is audible and obvious breathing that
impacts the conversation, such as a gasp, label it accordingly with the corresponding speech
label and following the appropriate rules.
A segment end should include audible non-verbal sounds (such as breath or continuation of
sibilant 's' sounds), and that they should be part of the same segment.
Any vocalization that conveys meaning, functions as feedback or acknowledgment, or is clearly
tied to the speaker's intent or delivery (e.g., sighs, gasps, heavy inhales, "sarcastic" throat
sounds) should be labeled as part of speech.
6. Sounds not to label
Do not label any of the following human sounds:
a. Speech from crowds: an audience laughing or cheering:
b. Crying:
c. Coughing:
If more than two people are simultaneously talking or making verbal sounds (laughter, "uh-huh,"
"mmmm," and so on), accurate labeling can be very difficult and time-consuming. Just label
what you can, but do not spend too much time on it.
7. Segmenting rules for pauses and intros
7.1 Do not create multiple segments if an intro, ad, and so on (which should be marked as
"Ignore") contains multiple voices. Create a single segment.
7.2 If a single speaker takes a pause longer than 2 seconds and then resumes, mark the
segment before and segment after into their own segments.
A less than 2-second pause should not result in the creation of a new segment.
8. Rules for ignoring content
8.1 To decide if a voice should be ignored use the context present in the audio. For example, if a
podcast presenter says a commercial will follow, and then you hear what sounds like a
commercial, select Ignore
8.2 Also, select Ignore for any opening segment or clip that comes before the main content of
the provided audio content (for example, a podcast) if it includes outside speakers (meaning a
speaker other than the podcast host, for example).
Such segments often arise at the beginning of a podcast (or TV or radio show) episode, before
the host actually starts speaking to introduce the episode.
In these cases, if context supports that it is the podcast host speaking, or if the only audio is the
podcast host giving a monologue that could be an advertisement or sponsorship, do not select
Ignore. Give a Speaker label in these instances.
These cases may be ambiguous and require you to use your best judgement to determine if it is
the podcast host, or an external speaker from an outside production.
8.3 Ignore pre-recorded segments (for example, sports event-could contain player voices, crowd
cheering-and commentators-more than one), that are not a part of the conversation.
Such segments may start in the foreground, could pass into the background (for example, when
podcast hosts talk over the recording), and then stop or fade away -
the key is whether or not the speakers have been arranged to be part of the dialogue with the
primary speaker or speakers, or not.
The guideline to ignore does not apply to all pre-recorded segments. If the pre-recorded
segments are interleaved somewhat naturally as part of the foreground conversation (for
instance, a prerecorded interview on a topic discussed in the audio) segment and label as if the
pre-recorded segments are part of the conversation.
9. Echoing voices
When they are part of the speaker's original turn-should be labeled as part of the same
segment, with the end time placed at the final audible echo rather than the end of the spoken
word.
E.g., in "Wondrium... um...um", where the two last "um" pieces are echoes, the end of the
segment should be on the last "um", not on the end of the initial word, "Wondrium".
Example