Fusion of Sentence-BERT and Machine Learning for Comment Text Topic Identification

Yuhan Wang, Bangguo Tong · 2024

In order to solve the problems of insufficient semantic description and poor semantic coherence of learned topics in comment text topic recognition. In this paper, Sentence-BERT sentence embedding model and LDA model are combined to improve the semanticity of comment text topics. The Sentence-BERT model is used to obtain the vector features at the sentence level of the comment text, while the LDA model is used to obtain the probabilistic topic vectors of the comment text, followed by connecting the two sets of vectors using an autoencoder, and applying the K-means algorithm to cluster the potential space vectors, and obtaining the contextual topic information from the class clusters. Through the experiments on the comment text dataset, the method of this paper can better obtain the topic words with semantic information. The fusion of Sentence-BERT model and LDA increases the complexity of the model. By comparison, the topic consistency index (Coherence) obtained by this paper's method is better than the current common comment text topic recognition methods.

Read the paper · More papers on PaperTik