Bayesian class-based language models
Yi Su · 2011
By capturing the intuition of "similar words appear in similar context", the Class-based Language Model (CLM) has found success from research projects to business products. How ever, most CLMs make a simplifying assumption that one word belongs to one class, which models poorly the fact that many words have multiple senses thus should belong to multiple classes. We propose a Bayesian formulation of the CLM, where a many-to-many mapping between words and classes, i.e., soft clustering, are naturally supported. A simple collapsed Gibbs sampler is provided to carry out the inference. Not only did we achieve a 22% relative reduction in perplexity on a Wall Street Journal corpus, but also reduced the word error rate of a state-of-the-art conversational telephony speech recognizer by 6% relative.