Robust speech recognition using maximum likelihood neural networks and continuous density hidden Markov models
Dong-Suk Yuk, ChiWei Che, James L. Flanagan · 2002
The laboratory performance of well trained speech recognizers is usually degraded when they are used in real world environments. Robust speech recognition is therefore an important issue for successful application of speech recognizers. Neural network based transformation methods are studied to compensate for the mismatched conditions of training and testing. First, a feature transformation neural network is studied. Second, a maximum likelihood neural network is applied to model transformations. The advantage of the neural network based transformation methods is that retraining of the speech recognizer for each particular environment is avoided. Furthermore, because the multi layer neural network is known to be able to compute nonlinear functions, the neural network based transformation methods are able to establish nonlinear mapping functions between training and testing environments without specific knowledge about the distortion or the mismatched environments.