Speech Separation Based on The Statistics of Binaural Auditory Features
Guy J. Brown, S. Harding, Jon Barker · 2006
A computational auditory scene analysis (CASA) system is described, in which sound separation according to spatial location is combined with the 'missing data' approach for automatic speech recognition. Time-frequency masks for the missing data recognizer are derived from the statistics of interaural time and level differences; these masks identify acoustic features that constitute reliable evidence of the target speech signal. It is demonstrated that this approach yields good performance in a challenging environment, in which a target voice is contaminated by another talker and reverberation. The ability of the system to generalize to source-receiver configurations that were not encountered during training is discussed