Match More, Extract Better! Hybrid Matching Model for Open Domain Web Keyphrase Extraction

Mingyang Song, Liping Jing, Yi Wei Feng · 2024

Keyphrase extraction aims to automatically extract salient phrases representing the critical information in the source document.Identifying salient phrases is challenging because there is a lot of noisy information in the document, leading to wrong extraction.To address this issue, in this paper, we propose a hybrid matching model for keyphrase extraction, which combines representation-focused and interactionbased matching modules into a unified framework for improving the performance of the keyphrase extraction task.Specifically, Hybrid-Match comprises (1) a PLM-based Siamese encoder component that represents both candidate phrases and documents, (2) an interactionfocused matching (IM) component that estimates word matches between candidate phrases and the corresponding document at the word level, and (3) a representation-focused matching (RM) component captures context-aware semantic relatedness of each candidate keyphrase at the phrase level.Extensive experimental results on the OpenKP dataset demonstrate that the performance of the proposed model Hybrid-Match outperforms the recent state-of-the-art keyphrase extraction baselines.Furthermore, we discuss the performance of large language models in keyphrase extraction based on recent studies and our experiments.

Read the paper · More papers on PaperTik