Dealing with noise in automatic speech recognition.
Douglas D. O’Shaughnessy · The Journal of the Acoustical Society of America · 2010
While automatic speech recognition (ASR) can work very well for clean speech, recognition accuracy often degrades significantly when the speech signal is subject to corruption, as occurs in many communication channels. This paper will survey recent methods for handling various distortions in practical ASR. The problem is often presented as an issue of mismatch between the models that are created during prior training phases and unforeseen environmental acoustic conditions that occur during the normal test phase. As one can never anticipate all possible future conditions, ASR analysis must be able to adapt to a wide variety of distortions. Human listeners furnish a useful standard of comparison for ASR in that humans are much more flexible in handling unexpected acoustic distortions than current ASR is. Methods that adapt ASR features and models will be compared against ASR methods that enhance the noisy input speech. Other topics to be discussed will include estimation of noise and channel parameters, RASTA, and cepstral mean normalization. TRAP-TANDEM features Vector Taylor Series, joint speech and noise modeling, and advanced front-end feature extraction. Single-microphone versus multi-microphone approaches will also be discussed.