Domain-independent term extraction & term network for scientific publications

Zheng Chen, Erjia Yan · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 2017

Term extraction is an essential tool for content-based publication analysis, and has a long history dating back to 1970s. However, previous methods are either domain-specific, or need complex model training, or relies on external resources like Wikipedia. Recent rise of cross-domain publication content analyses has put forward the demand for simple and efficient domain-independent extraction method. This paper proposes a new rule-based method that adapts C-value method to publication analysis, extends it with two types of frequency lists and sigmoid functions, and develops a prototype term extraction system. Our experiment shows a remarkable reduction of “error” with better or competitive “keyword recall” against C-Value method and a complex term extraction method provided by Translated.net. We then construct a term network by connecting adjacent terms in a paragraph and demonstrate that rich and meaningful analysis can be done on such network through a case study on an HCI abstract corpus.

Read the paper · More papers on PaperTik