Keyphrases extraction: Approach Based on Document Paragraph Weights
Lahbib Ajallouda, Ahmed Zellou, Imane Ettahiri, Karim Doumi · 2022
In recent years, the exploitation of sentence embedding techniques in natural language processing field has encouraged the proposal of new methods for extracting keyphrases from documents based on these techniques. Most of these approaches select keyphrases from a set of candidate phrases based on their semantic proximity to the document. In general, most documents contain complementary paragraphs that are unrelated to the topics covered. This factor reduces the credibility of the semantic proximity of candidate keyphrases to the document. Exploitation of document paragraphs weights during the semantic similarity calculation will inevitably improve the performance of keyphrase extraction from document. In this paper, we propose a new method to extract keyphrases based on document paragraphs weights. Our method is based on sentence embedding techniques and semantic proximity of candidate key phrases from document paragraphs. We evaluated the proposed method on three datasets, Inspec, Semeval2010 and KPTimes, where our results showed that that using document paragraph weight improved the performance of keyphrases extraction.