Experimental Study of Speech Recognition in Noisy Environments

Tomáš Kreisinger, Pavel Sovka, Petr Pollák, Jan Uhlíř · Applied and numerical harmonic analysis · 1998

Achieving reliable performance in a speech recognizer for car telephone applications has been studied intensively for more than a decade. This paper addresses the effects of mismatched conditions and their minimization with respect to the performance of speaker-independent isolated-word recognition in a car-noise environment without considering the Lombard effect. This study is primarily intended to evaluate the dependence of the recognition rate on the signal-to-noise ratio (SNR) of an input signal either without any noise-compensation method or with a noise-compensation or noise-adaptive method and especially to find the appropriate conditions so that an isolated word recognizer can be used in a real car-noise environment. When hidden Markov models (HMMs) are trained on noisy speech with a SNR of l0dB, it is possible to recognize noisy speech with a SNR in the interval from 40 dB to 5 dB with a recognition rate better than 93%. If modified spectral subtraction is used and models are trained on the enhanced speech, the SNR interval increases to 0 dB. If the parallel model combination (PMC) technique is used, there is no need to train models on noisy or enhanced speech. The model adaptation enables recognizing noisy speech with any SNR from 40 to -10 dB with a recognition rate greater than 73% (for a SNR from 40 to 5 dB, the recognition rate is above 93%). In this respect PMC offers great flexibility with better recognition rates than other noise-compensation techniques.

Read the paper · More papers on PaperTik