Toward better automatic speech recognition
Douglas D. O’Shaughnessy, W. Wang, Wen Zhu, Vincent Barreaud, T. Nagarajan, Rangarao Muralishankar · Canadian acoustics · 2005
Automatic speech recognition (ASR) performs best when there is a strong correspondence between system training and operating conditions, e.g., when one tests on speech data that is similar in style to that used for training.Mismatch (different environment speakers, or vocabulary) degrades performance.We have developed new model techniques able to adapt to various speech environments without modifying the basic ASR systems: an appropriate feature transformation scheme for the Mel-frequency cepstral coefficients (MFCC), a new speech-processing front-end feature that performs better than the existing MFCC, a log-energy dynamic range normalization technique for ASR in adverse conditions, and a continuous ASR method that exploits the advantages o f syllable and phoneme-based sub-word unit models.