A Keyword Extraction Method Integrating Word Semantic Similarity and Topic Model
Weifeng Liu, Wanyu Li, Liwen Ma · 2024
The existing unsupervised random walk keyphrase extraction model mainly relies on the global structural information of the word graph. It ignores the external structure and topic information of the document. To address this problem, this paper proposes a keyphrase extraction method that integrates word sense similarity and document implicit topic information. First, the latent Dirichlet allocation (LDA) topic model is used to calculate the implicit topic distribution in the document. It can quantify the preference of each word for different topics in the document. The random jump probability in the graph is converted into the importance of the word on the implicit topic. Secondly, the word sense similarity is introduced to reset the influence between different words. The importance weights of the words in the original word graph are converted into non-uniform transfer. Finally, the candidate keyword weights are calculated through iteration. The selected keywords are expanded into valid keyphrases based on the grammar. The TopN candi-date keyphrases with higher importance scores are selected as document keyphrases. The improved algorithm performs better than TextRank and SingleRank. The Fl-Scores are improved by 8.51 % and 4.92 % respectively. Experimental results show that this method can make full use of document topic information and word sense similarity. It makes the automatic keyphrase extraction effect more significant.