A comparison of distance measures for clustering in speaker diarization

Marcelo de Campos Niero, Álvaro Veiga, André Adami · 2014

Speaker diarization consists in answering the question “Who spoke when” for a given conversation in a telephone call, meeting, or broadcast news, without any prior information about neither the audio nor the speakers. Speaker diarization task emerged as a way to optimize audio information retrieval processing by detecting and tracking speech and speaker information. Computationally speaking, the diarization processing occurs through four main steps: feature extraction of signal, speech and non-speech detection, segmentation and clustering. In this work, the clustering step is analyzed by comparing distance measures commonly used in current speaker diarization systems. The results show that pairs of clusters with a large difference in the number of data samples are more sensitive to errors, the number of mixtures of an external model affects the discriminative power of distance measures, and the number of estimated parameters affects the speaker discrimination. All experiments are performed on an excerpt from TIMIT corpus and the diarization task database used in the 2002 NIST Speaker Recognition Evaluation.

Read the paper · More papers on PaperTik