Semantic Class Induction for Language Model Adaptation in a Chinese Voice Search System

Yali Li, Weiqun Xu, Changchun Bao, Li Ta, Jielin Pan, Yonghong Yan · 2010

In this paper we describe our work on generating in-domain corpus using auto-induced semantic classes and structures for language model adaptation in a voice search dialogue system. We proposed a novel similarity measure based on co-occurrence probabilities for inducing semantic classes. Clustering with the new similarity measure outperformed that with the widely used distance measure based on Kullback-Leibler divergence. For language model adaptation, we adopted the widely used approach of model interpolation. Experiments show that both human-human and generated data helped a lot and the latter helped more. This means that the generated data is more in-domain than the human-human data for human-computer dialogues. The performance of 9.0% in character recognition error rate and 25.5 in perplexity on the test data is achieved with a language model from an interpolated language model.

Read the paper · More papers on PaperTik