A NOVEL WORD CLUSTERING ALGORITHM AND VARI GRAM LANGUAGE MODEL
Lang Chen · Chinese Journal of Computers · 1999
Class based statistic language model is an important method to solve the problem of sparse data. But there are two bottlenecks about this model: (1) The problem of word clustring, it is hard to find a suitable clustering method that has good performance and not large amount of computation. (2) Class based method always lose some prediction ability to adapt the text of different domain. The authors try to solve above problems in this paper. This paper presents a hierarchical word clustering algorithm based on the similarity between words in nature language. Experiments show that this method is better than conventional greedy clustering method in speed and performance. At the same time, this paper presents a new method to create the vari gram model, and gets a class based vari gram model that performance is much better than traditional class based n gram model.