Detecting the Number of Speakers in Speech Mixtures by Human and Machine

Tomasz Mąka, Mirosław Łazoryszczak · 2018

The problem of sound sources estimation and its properties in acoustic scene plays important role in many voice-based interaction systems. The interference between sources can deteriorate system performance meaningfully. The paper presents a comparison results of objective and subjective methods applied to the process of identification the number of speakers in speech mixtures. The audio data set used for computational and subjective tests consists of a number of utterances spoken by from two up to seven simultaneous speakers. In order to determine the number of speakers, two approaches are applied to speech mixtures: first uses spectrogram factorization with NMF (non-negative matrix factorization) algorithm, the other is based on the perceptual evaluation by the group of listeners. Both techniques are compared in terms of classification accuracy.

Read the paper · More papers on PaperTik