Centrality Measures for Non-Contextual Graph-Based Unsupervised Single Document Keyword Extraction
Natalie Schluter · 2014
The manner in which keywords fulfill the role of being central to a document is frustratingly still an open question. In this paper, we hope to shed some light on the essence of keywords in scientific articles and thereby motivate the graph-based approach to keyword extraction. We identify the document model captured by the text graph generated as input to a number of centrality metrics, and overview what these metrics say about keywords. In doing so, we achieve state-of-the-art results in unsupervised non-contextual single document keyword extraction. In this paper, we present a study on the essence of keywords in scientific articles and thereby motivate the graph-based approach to keyword extraction. We identify the document model captured by the text graph generated as input to a number of centrality metrics, and overview what these metrics say about keywords. In doing so, we achieve state-of- the-art results in non-contextual single document keyword extraction for the Inspec corpus, an NDCG score of 0.07578 ; we can affirm this, because the systems we compare here in terms of NDCG include re-implementations of the previous state-of-the-art. We identify two broad types of single document keyword extraction (SDKE). Contextual SDKE makes use of the document set to which the relevant document belongs, and in which there are similar documents ; other information outside of the document set may also be used in some types of contextual SDKE. Non-contextual SDKE makes use of only the relevant document with no other information. The latter does not necessarily make the assumption of independence of documents in general. Non-contextual SDKE is actually important for the case of isolated documents (not part of a document set), as well as for documents for which relevant supplementary information may be non-existent or unreliable. In addition, the study of non-contextual SDKE is a study of baselines in keyword extraction : before constructing complex methods using more information, it is important to understand what we can achieve with the document alone.