PENS: a confidence parameter estimating the number of speakers

Siham Ouamour, Mhania Guerti, Halim Sayoud · ExLing Conferences · 2008

Is it possible to know how many speakers are speaking simultaneously in a case of speech overlap? While the human brain—a creation not yet mastered—manages to do this and even to understand the meaning of the mixed speech, it is not yet the case for existing automatic systems. For this task, we propose a new method able to estimate the number of speakers in a mixture of speech signals. The algorithm developed here is based on the computation of the statistical characteristics of the seventh Mel coefficient extracted by spectral analysis from the speech signal. This algorithm, which uses a confidence parameter that we called PENS, is tested on seven different sets of the ORATOR database, each containing seven multi-speaker files. Results show that the PENS parameter permits a clear discrimination, without any ambiguity, between a mono-speaker signal (only one speaker is speaking) and a mixed-speaker signal (several speakers are speaking simultaneously). Moreover, in the case of mixed speech signals, it permits an estimation of the number of speakers with good precision, especially when the number of speakers is less than four.

Read the paper · More papers on PaperTik