Detection, separation and recognition of speech from continuous signals using spectral factorisation
Antti Hurmalainen, Jort Florent Gemmeke, Tuomas I. Virtanen · Lirias · 2012
In real world speech processing, the signals are often con-tinuous and consist of momentary segments of speech over non-stationary background noise. It has been demonstrated that spectral factorisation using multi-frame atoms can be suc-cessfully employed to separate and recognise speech in ad-verse conditions. While in previous work full knowledge of utterance endpointing and speaker identity was used for noise modelling and speech recognition, this study proposes spec-tral factorisation and sparse classification techniques to de-tect, identify, separate and recognise speech from a continu-ous noisy input. Speech models are trained beforehand, but noise models are acquired adaptively from the input by us-ing voice activity detection without prior knowledge of noise-only locations. The results are evaluated on the CHiME cor-pus, containing utterances from 34 speakers over highly non-stationary multi-source noise. Index Terms — Spectral factorization, speech recogni-tion, speaker recognition, voice activity detection, speech separation 1.