Automatic Clustering of Languages Based on Probabilistic Models

Kenji Kita · Journal of Quantitative Linguistics · 1999

This paper proposes a novel method for automatically clustering languages. The basic idea of this method involves developing a probabilistic model for each language from the given linguistic data, and then computing the distances between languages according to the distance measure defined on the language models. Clustering is performed based on this distance measure. The effectiveness of the proposed method has been confirmed by evaluation experiments using two kinds of data: (1) word lists of sixteen languages obtained from the Oxford Text Archive, and (2) multilingual texts of nineteen languages from the ECI/MC1 corpus.

Read the paper · More papers on PaperTik