Speech Recognition in Unknown Noisy Conditions
Ji Ming, Baochun Hou · 2007
Robust Speech Recognition and Understanding 176 severely affected by noise can thus be exploited for recognition.This assumption is not realistic for many real-world applications in which the noise will affect all time-frequency components of the speech signal, i.e., we face a full feature corruption problem.In this chapter, we investigate speech recognition in noisy environments assuming a highly unfavourable scenario: an accurate estimation of the nature and characteristics of the noise is difficult, if not impossible.As such, traditional techniques for noise removal or compensation, which usually assume a prior knowledge of the noise, become inapplicable.We describe a new noise compensation approach, namely universal compensation, as a solution to the problem.The new approach combines subband modeling, multicondition model training and missing-feature theory as a means of minimizing the requirement for the information of the noise, while allowing any corruption type, including full feature corruption, to be modelled.Subband features are used instead of conventional fullband features to isolate noisy frequency bands from usable frequency bands; multicondition training provides compensations for expected or generic noise; and missing-feature theory is applied to deal with the remaining training and testing mismatch, by ignoring the mismatched subbands from scoring.The rest of the chapter is organized as follows.Section 2 introduces the universal compensation approach and the algorithms for incorporating the approach into a hidden Markov model for speech recognition.Section 3 describes experimental evaluation on the Aurora 2 and 3 tasks for speech recognition involving a variety of simulated and realistic noises, including new noise types not seen in the original databases.Section 4 presents a summary along with the on-going work for further developing the technique.