Including uncertainty of speech observations in robust speech recognition

Jose Carlos Segura, Ángel de la Torre, Javier Ramı́rez, Antonio J. Rubio, Carmen Benı́tez · 2004

Noise compensation methods for speech recognition provide a cleaned version of the speech representation. Usually this cleaned version is the expected value of the speech parameters given the observed noisy speech and the noise statistic. A more realistic representation should include the probability distribution of the cleaned speech instead of its expected value in order to represent the uncertainty associated to the compensation process due to the variability of the noise process. Recently, the inclusion of the uncertainty in the recognition process has been studied. Some approaches represent the uncertainty in the HMM parameters values. Other approaches represent it in the feature space. This second approach offers a much simpler system implementation and lower computational cost. In this paper we have developed a noise compensation technique that incorporates the variance of the cleaned speech into the speech representation. The variance is estimated using a Wiener filter during the speech feature enhancement process. This way of including the uncertainty implies the modification of the decoding rule. Experimental results using AURORA 2 database demonstrate a sustained improvement of the performance in the recognition system (about 21% word error rate reduction) when uncertainty is considered in the decoding rule.

Read the paper · More papers on PaperTik