Speaker Clustering of Stereo Audio Documents Based on Sequential Gathering Process

Halim Sayoud, Siham Ouamour · J. Inf. Hiding Multim. Signal Process. · 2010

This paper focuses on the use of sequential speaker clustering of stereo audio documents to obtain a classification of the different speech segments contained in those documents, according to the speakers who are participating in the audio recording. In general, speaker clustering is used as a second step in a global system of speaker diariza- tion, where the first step deals with the task of speaker segmentation. However, in some applications, the term speaker diarization is confused with speaker clustering. In such applications, the homogeneous segments are automatically separated, like in telephone answering machines or vocal boxes, where the vocal messages are already separated by a sound beep. Even though in our global project we use two main techniques based on speaker localization and speaker discrimination, in this paper we will describe only the second technique, which uses a sequential clustering approach, in order to gather the similar homogeneous segments into classes of speakers. Each class contains the global intervention of only one speaker in the entire audio document. The sequential cluster- ing approach uses a mono-gaussian measure ( G) that allows us to assess the degree of similarity between the different homogeneous segments. The application concerns the clustering of stereo debates between several speakers who are located at different positions in the meeting-room. For the evaluation, experiments are conducted on a stereophonic database called DB15, which is composed of 15 scenarios of about 3.5mn each and con- taining two or three speakers speaking sequentially in every scenario. The new algorithm shows good performances, when the length of the speech segments is over 4s. Keywords: speaker clustering, speaker diarization, sequential clustering, stereo audio document, sec- ond order statistical measures, mono-gaussian measures.

Read the paper · More papers on PaperTik