Overlapped Speech Detection and Competing Speaker Counting–‐Humans Versus Deep Learning
Valentin Andrei, Horia Cucu, Corneliu Burileanu · IEEE Journal of Selected Topics in Signal Processing · 2019
A natural evolution of applications that analyze speech is to improve their robustness to multi-speaker environments. Humans use selective auditory attention and can easily switch focus from one source to another even when listening to a single channel recording with overlapped speech. The same brain feature allows us to detect the number of simultaneously active sources. In order to quantify human level performance for this task we have designed a perception study that evaluates participants' ability to accurately count multiple speakers in a single channel audio file. We also analyzed the influence of listening time and of hearing familiar voices. The study was carried in 3 sessions, with the help of 31, 38 and 80 volunteers and significantly extends the findings in existing literature. Using the conclusions from the perception analysis, several convolutional neural networks were trained to estimate the number of competing speakers on speech timeframes ranging from 25 ms to 1 s. The same models were instructed to tag overlapped speech and we observed F-score values up to 0.91. For both tasks, the proposed methods lead to lower error than existing approaches and require smaller timeframes. Compared with human listeners, the neural networks can count speakers more accurately by analyzing a considerably shorter recording.