Improving automatic speech recognition via better analysis and adaptation
Douglas D. O’Shaughnessy, Wayne Wang, William Zhu, Vincent Barreaud, T. Nagarajan, Rangarao Muralishankar · The Journal of the Acoustical Society of America · 2005
One way to improve automatic speech recognition (ASR) systems is to reduce the mismatch between system training and operating conditions, as such mismatch seriously degrades performance. We have developed model adaptation techniques able to adapt to various speech environments without modifying ASR systems, and have developed an appropriate feature transformation scheme for the Mel-frequency cepstral coefficients (MFCC), a popular front-end feature of ASR systems. We use maximum a posteriori model adaptation and a method based on Bayesian parametric representation. Feature transformation aims to maximize the desired source of information for a given speech signal in the front-end features and to minimize undesired sources. Frequency-domain autoregressive modeling and a segmentation algorithm are being developed, e.g., to segment a speech signal into syllablelike units. We also introduce a new speech-processing front-end feature that performs better than the existing MFCC, as well as a log-energy dynamic range normalization technique for ASR in adverse conditions. In addition, we have developed a continuous ASR method that exploits the advantages of syllable and phoneme-based subword unit models. [Work supported by NSERC-Canada and Prompt-Quebec.]