Speaker Diarization
Partitioning an audio recording into speaker-homogeneous regions and assigning consistent anonymous speaker labels over time.
Who Spoke When?
Speaker diarization answers a timeline question: which portions of the recording belong to the same speaker. It does not necessarily identify the real-world identity of that person.
A common pipeline includes speech activity detection, segmentation, speaker embeddings and clustering, followed by optional resegmentation.
Main Failure Modes
Overlapping speech, short turns, channel changes, background noise and incorrect VAD boundaries can all degrade clustering. Diarization error therefore comes from more than the embedding model.
Deterministic stereo cases can sometimes avoid generic clustering when each channel has known semantic meaning; I discuss that distinction in Deterministic Speaker Separation in Stereo Streams.
Related Concepts
Related article: Audio Segmentation
Related article: Speaker Recognition and Diarization
Direct source: The primary paper or official specification for Speaker Diarization is linked here for verification.