An Environment Compensated Maximum Likelihood Training Approach Based on Stochastic Vector Mapping
Jian Wu, Qiang Huo, Donglai Zhu · 2006
Several recent approaches for robust speech recognition are developed based on the concept of stochastic vector mapping (SVM) that perform a frame-dependent bias removal to compensate for environmental variabilities in both training and recognition stages. Some of them require stereo recordings of both clean and noisy speech for the estimation of SVM function parameters. In this paper, we present a detailed formulation of a maximum likelihood training approach for the joint design of SVM function parameters and HMM parameters of a speech recognizer that does not rely on the availability of stereo training data. Its learning behavior and effectiveness is demonstrated by using the experimental results on the Aurora3 Finnish connected digits database recorded by using both close-talking and hands-free microphones in cars.