Comparative Analysis of Embedding Models for Keyphrase Extraction: A KeyBERT-Based Approach

Nimisha MR, T Shamitha, Geetishree Mishra · 2023

Key phrase extraction is a fundamental task in information management, which is often used as a preliminary step in various information retrieval and natural language processing tasks. Embedding models achieve higher F-scores than graph-based models in key phrase extraction for short documents compared to longer documents. Therefore, they are suitable for real-time semantic processing of short textual data. With a vast set of embedding models dedicated to key phrase extraction, it becomes a tedious task to overview and compare the performance of each model. Hence, we have carried out an extensive comparative study of four prominent pre-trained embedding models, namely Sentence Transformers, Flair, spaCy, and Universal Sentence Encoder, that have been used in conjunction with Key-BERT. The results of the experiments conducted on this corpus show that the Sentence Transformer “all-MiniLM-L6-v2“ version outperforms the other embedding models listed above, with respect to BERT score evaluation and adaptability to transfer learning. Notably, spaCy exhibits fast processing while maintaining satisfactory accuracy. An extensive comparison of Key-BERT-based embedding models suited to texts of different sizes and a detailed qualitative analysis makes the case for our proposed methodology.

Read the paper · More papers on PaperTik