Robust Automatic Speech Recognition with Missing and Unreliable Data
Ljubomir Josifovski · 2003
Automatic speech recognition (ASR) systems have made dramatic performance leaps in the recent past. Yet, the notion that the key to making recognition more robust is to reduce the di#erence between training and test conditions is still commonly held. As ASR applications move from tightly controlled to more natural environments with a varying number of unpredictable sound sources, this assumption is becoming less and less viable. Decoding the speech source of interest while listening to several sound sources at the same time seems a more accurate description of the ASR process that suits these challenging environments. This thesis discusses the theoretical and practical issues which arise from this viewpoint. The aim is to explore the division of the problem of robust ASR into two subproblems: (a) identification/separation of the speech and noise using speech properties alone; and (b) recognition based on the resulting partial evidence. The basic assumption is that some regions of the speech time-frequency representation remain relatively una#ected by the noise, that they can be identified and that they alone are su#cient for ASR. In contrast to conventional techniques which require models of all sources in the auditory scene and their subsequent decoding even when only one of the sources is of interest, the techniques described in this thesis make no such requirement. However, they are flexible enough to use this information if it is available.