Keyphrase Extraction in Russian and English Scientific Articles Using Sentence Embeddings

Quang Huy Nguyen, Mark Zaslavskiy · 2021

Keyphrases provide an overview of the articles, making it a powerful tool for categorizing scientific articles. This paper introduces and describes our supervised machine learning model for automatic keywords extraction. The model calculates features from traditional statistical metrics and new state-of-the-art sentence embeddings to predict a confidence score annotating conformity of keyphrase candidate. The model is tested on corpora of Russian as well as English scientific articles. When compared to the chosen baseline methods of the experiment, our model achieved a comparable F1 score when applied to the Russian corpora; and outperformed them when applied to the English corpora. Using F1-score as the evaluation metric, we also experimented with the model's parameters, such as the embedder and the set of features used as input. We found the pre-trained embedder that provides the best possible outcome for our task and confirmed that our model works best with the full set of features - non of the input to the model is redundant. For future works, we set our goal on deploying the model on existing system. Moreover, we suggest training a delicated embedding module to improve the model performance when working with articles written in Russian.

Read the paper · More papers on PaperTik