A Word Clustering Method Based on Mutual Information
Yuan Li-chi · Systems Engineering · 2008
Cluster-based statistic language model is an important method for solving the problem of sparse data.Conventional statistical clustering methods usually base on greedy principle.The common standard for evaluating a clustering algorithm is the likelihood function or perplexity of the corpus.Conventional clustering algorithms often converge to a local optimum,so global optimum is not guaranteed,and initial choices can influence the final results.In order to solve these problems,we first give a definition to word similarity by utilizing mutual information,then,based on which,define word set similarity,and finally propose a bottom-up hierarchical clustering algorithm based on similarity.This method can not only improve clustering effect,but also choose different definitions of similarity for different cluster-based models,such as predictive clustering,conditional clustering,and combined clustering,thus improving the effect of using clusters.Experiments show that word clustering algorithm based on similarity is better than conventional greedy clustering method in terms of calculating complexity and clustering effect.