Keyphrase Extraction Combining Word and Document Embeddings
ZU Xian, XIE Fei, LIU Xiaojian · DOAJ (DOAJ: Directory of Open Access Journals) · 2021
With the increasing amount of text data in various application fields, how to quickly and accurately extract the main information has become the main task of keyphrase extraction. This paper proposes a novel method for keyphrase extraction based on word and document vectors. By calculating vector representation between word and document on the same dimensional vector space, the semantic similarity between word and document can be got, which can be used as the initial weight of each word node in the undirected graph. Then, this paper calculates the score of each word and candidate phrase using a semantic biased random walk strategy. Finally, the top[N] scored candidate phrases are selected as the final keyphrases. Experimental results on the public datasets show that the proposed algorithm outperforms the state-of-the-art keyphrase extraction methods in precision, recall, and F-measure. It can greatly improve the efficiency of automatic keyphrase extraction.