Evaluation of PEMO in robust speech recognition
Klaus Kasper, Herbert Reininger · The Journal of the Acoustical Society of America · 1999
A major problem in speech recognition is the robustness of recognition performance against additive background noise and convolutive distortions due to changing transmission channels or varying microphone characteristics. One approach to overcome this problem is to apply PEMO [Dau, Püschel, and Kohlrausch, J. Acoust. Soc. Am. 99, 3615–3622 (1996)] for extraction of noise robust feature vectors. The recognition performance achievable with PEMO in combination with a locally recurrent neural network (LRNN) or hidden Markov models (HMM) for feature scoring was evaluated in a task of speaker-independent word recognition. The robustness was tested and compared with that of other feature types by applying the speech recognition systems to speech signals disturbed by additive background noise and to speech signals recorded over telephone lines. It was found that only with LRNN it is possible to exploit the potential of PEMO, not with HMM. A detailed analysis of the interplay between PEMO features and the scoring techniques revealed that the long-term dependencies in the sequence PEMO features can only be captured by LRNN. Furthermore, LRNN discriminate between the distinct and sparse peaks of the PEMO speech representation—which are well maintained also in the case of noisy speech—and artifacts introduced by the additive or convolutive noise.