Short Text Clustering Algorithm Based on Frequent Closed Word Sets

Chunxia Jin, Qiuchan Bai · 2019

The text mining of micro-blog topic information can effectively obtain the attention degree of internet users for news events. It is of great significance in the field of public opinion monitoring and analysis. At the situation of the algorithm of traditional frequent word set is suitable for long text information clustering, this paper proposes to mine top-K frequent corpus in short text database and then to divide micro-blog topic texts covering the same frequent word sets into the same cluster. Combined with the largest frequent word-sets for similarity calculation, the overlapped document is re-divided to achieve micro-blog short text clustering. The experimental results of micro-blog topic dataset and the comparison with K-means clustering algorithm show that the proposed algorithm can effectively solve the sparseness and high-dimension problem of micro-blog topic short text clustering and greatly improve the micro-blog short text clustering effect.

Read the paper · More papers on PaperTik