Research on Statistical Word Clustering Methods

Lichi Yuan · 2019

Statistic language model based on classes of words is an effective method to solve sparse-data problems. Conventional statistical word clustering models usually base on greedy principle. The common Metric for evaluating a clustering model is the perplexity of the corpus or likelihood function. Conventional statistical word clustering algorithms often converge to a local optimum, so global optimum is not guaranteed, and initial choices can influence the final clustering result. Pointing to these problems, this paper presents a definition of word similarity by utilizing mutual information, and gives the definition of word set similarity based on word similarity, and puts forward a bottom-up hierarchical word clustering algorithm which can get global optimum. Experimental results show that the word clustering algorithm is of high executing speed and has good clustering performance. We then interpolated the class-based models with the word-based models and found that it mitigates the remaining sparse-data problems.

Read the paper · More papers on PaperTik