Automatic speech recognition system channel modeling
Qun Feng Tan, Kartik Audhkhasi, Panayiotis Georgiou, Emil Ettelaie, Shrikanth Shri Narayanan · 2010
In this paper, we present a systems approach for channel mod-eling of an Automatic Speech Recognition (ASR) system. This can have implications in improving speech recognition com-ponents, such as through discriminative language modeling. We simulate the ASR corruption using a phrase-based machine translation system trained between the reference phoneme and output phoneme sequences of a real ASR. We demonstrate that local optimization on the quality of phoneme-to-phoneme map-pings does not directly translate to overall improvement of the entire model. However, we are still able to capitalize on contex-tual information of the phonemes which a simple acoustic dis-tance model is not able to accomplish. Hence we show that the use of longer context results in a significantly improved model of the ASR channel.