Algorithms for Joint Evaluation of Multiple Speech Patterns for Automatic Speech Recognition
Nishanth Ulhas, T.V. Sreenivas · InTech eBooks · 2008
Improving speech recognition performance in the presence of noise and interference continues to be a challenging problem.Automatic Speech Recognition (ASR) systems work well when the test and training conditions match.In real world environments there is often a mismatch between testing and training conditions.Various factors like additive noise, acoustic echo, and speaker accent, affect the speech recognition performance.Since ASR is a statistical pattern recognition problem, if the test patterns are unlike anything used to train the models, errors are bound to occur, due to feature vector mismatch.Various approaches to robustness have been proposed in the ASR literature contributing to mainly two topics: (i) reducing the variability in the feature vectors or (ii) modify the statistical model parameters to suit the noisy condition.While some of the techniques are quite effective, we would like to examine robustness from a different perspective.Considering the analogy of human communication over telephones, it is quite common to ask the person speaking to us, to repeat certain portions of their speech, because we don't understand it.This happens more often in the presence of background noise where the intelligibility of speech is affected significantly.Although exact nature of how humans decode multiple repetitions of speech is not known, it is quite possible that we use the combined knowledge of the multiple utterances and decode the unclear part of speech.Majority of ASR algorithms do not address this issue, except in very specific issues such as pronunciation modeling.We recognize that under very high noise conditions or bursty error channels, such as in packet communication where packets get dropped, it would be beneficial to take the approach of repeated utterances for robust ASR.We have formulated a set of algorithms for both joint evaluation/decoding for recognizing noisy test utterances as well as utilize the same formulation for selective training of Hidden Markov Models (HMMs), again for robust performance.Evaluating the algorithms on a speaker independent confusable word Isolated Word Recognition (IWR) task under noisy conditions has shown significant improvement in performance over the baseline systems which do not utilize such joint evaluation strategy.A simultaneous decoding algorithm using multiple utterances to derive one or more allophonic transcriptions for each word was proposed in [Wu & Gupta, 1999].The goal of a