Maximum Entropy Based Generic Filter for Language Model Adaptation
Dong Yu, M. Mahajan, Peter Mau, Alex Acero · 2006
Language model (LM) adaptation has been shown to be very important in reducing the word error rate (WER) in task specific speech recognition systems. Adaptation data collected in the real world, however, usually contain large amounts of non-dictated text, such as email headers, long URL, code fragments, included reply, signature, etc., that the user never dictates. Adapting with these data may corrupt the LM. We propose a maximum entropy (MaxEnt) based filter to remove a variety of non-dictated words from the adaptation data and improve the effectiveness of the LM adaptation. We argue that this generic filter is language independent and efficient. We describe the design of the filter, and show that the use of the filter can give us 10% relative WER reduction over LM adaptation without filtering, and 22% relative WER reduction over the unadapted LM in an English email dictation task.