Large-Scale Text Potential Topic Mining and Clustering Method Based on Natural Language Processing
Zuowei Chen, Ze-Wei Xia, Changyu Sun, Minjie Wang, Xuegang Xia · 2025
This study proposes a clustering method for large-scale text potential topics using natural language processing techniques. By addressing challenges such as data diversity, noise, and semantic complexity, the method constructs a Biterm Topic Model (BTM) to extract word-document-topic probability distributions, enabling effective clustering of text into potential topics. Using the Sogou Laboratory News Corpus and the Fudan University Text Classification Corpus, the proposed method demonstrates superior performance compared to benchmark clustering methods such as PV-DBOW and K-means. Experimental results show an average CH index improvement of 5.0% for the Sogou corpus and 4.8% for the Fudan corpus, highlighting the method's efficiency in semantic representation and clustering quality. This research provides valuable insights for text classification, topic identification, and information retrieval applications.