Research on Clustering Technology Based on Co-occurrence Word Model for Tibetan Microblog

Ailin Li, Tao Jiang, Qingshuai Wang, Hongzhi Yu · 2016

In order to solve the two key problems of the short text classification, very sparse features and strong context dependency, this paper proposed clustering technology based on Co-Occurrence Word Model for Tibetan microblog. The Tibetan news corpus (standard long text) as the experimental training corpus, using LDA model to construct the co-occurrence word network which related to one theme. According to the co-occurrence network relationship to determine the attribution of Tibetan microblog short text, so as to solve the problem of short text lack of semantic correlation and data sparseness, finally, we using K-means++ clustering algorithm to cluster. By comparing the results of three clustering methods, experimental results show that this method has obvious effect on the microblog Tibetan text clustering, and accuracy reached 88.06%.

Read the paper · More papers on PaperTik