Template mining for the extraction of citation from digital documents

Schubert Shou-Boon Foo, Gobinda G Chowdhury, Ying Ding · 2001

Information Extraction (IE) is a term that involves the activity of automatically extracting pre-specified sorts of information from short, natural language texts. It may be seen as the activity of populating a structured information source (or database) from an unstructured, or free text, information source. This structured database then can be used for a number of purposes: for creating a citation database, for report generation, for decision making in business, for using data-mining or artificial intelligent and neural network techniques, and so on. Template mining is a particular technique used in IE. When text matches a template, the system extracts data according to instructions associated with that template. This study hypothesizes that the template mining technique can be used to extract citation information from printed and digital full-text articles so that universal or semi-universal citation databases can be automatically established before too long in the future.

Read the paper · More papers on PaperTik