Web-Data Augmented Language Models for Mandarin Conversational Speech Recognition

T. Ng, Mari Ostendorf, Mei-Yuh Hwang, Man-Hung Siu, Ivan Bulyko, Xin Lei · 2006

Lack of data is a problem in training language models for conversational speech recognition, particularly for languages other than English. Experiments in English have successfully used Web-based text collection, targeted for a conversational style, to augment small sets of transcribed speech; we look at extending these techniques to Mandarin. In addition, we investigate different techniques for topic adaptation. Experiments in recognizing Mandarin telephone conversations show that the use of filtered Web data leads to a 28% reduction in perplexity and 7% reduction in character error rate, with most of the gain due to the general filtered Web data.

Read the paper · More papers on PaperTik