Discriminant training of front-end and acoustic modeling stages to heterogeneous acoustic environments for multi-stream automatic speech recognition
Michael L. Shire, Nelson H. Morgan · 2000
Automatic Speech Recognition (ASR) still poses a problem to researchers. In par-ticular, most ASR systems have not been able to fully handle adverse acoustic en-vironments. Although a large number of modications have resulted in increased levels of performance robustness, ASR systems still fall short of human recognition ability in a large number of environments. A possible shortcoming of the typical ASR system is the reliance on a single stream of front-end acoustic features and acous-tic modeling feature probabilities. A single front-end feature extraction algorithm may not be capable of maintaining robustness to arbitrary acoustic environments. Acoustic modeling will also degrade due to distributional changes caused by the acoustic environment. This thesis explores the parallel use of multiple front-end and acoustic modeling elements to improve upon this shortcoming. Each ASR acoustic modeling component is trained to estimate class posterior probabilities in a partic-ular acoustic environment. In addition to discriminative training of the probability estimator, existing feature extraction algorithms are modied in such a way as to