A Text Document Clustering Method Based on Weighted BERT Model

Yutong Li, Juanjuan Cai, Jingling Wang · 2020 IEEE 4th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC) · 2020

Traditional text document clustering methods represent documents with uncontextualized word embeddings and vector space model, which neglect the polysemy and the semantic relation between words. This paper presents a novel text document clustering method to deal with these problems. Firstly, pre-trained language representation model Bidirectional Encoder Representations from Transformers (BERT) is utilized to generate sentence embeddings. Then, two sentence-level weighting schemes based on named entity are designed to enhance the performance. Finally, the k-means clustering algorithm is applied to find groups of similar documents. Experimental results on four datasets indicate that the proposed weighted method achieves higher accuracy than unweighted average method. Friedman tests conducted separately with F1 score and Adjusted Rand Index (ARI) values both validate better overall performance of our proposed method.

Read the paper · More papers on PaperTik