Tag insertion complexity
Stuart Yeates, Ian H. Witten, David Bainbridge · 2002
This paper is about inferring markup information, a generalization of part-of-speech tagging. We use compression models based on a marked-up training corpus and apply them to fresh, unmarked, text. In effect, this technique builds filters that extract information from text in a way that is generalized because it depends on training text rather than preprogrammed heuristics. As illustrated, we use SGML tags to represent the extracted information. However, we work in a more controlled textual environment: we use bibliographic text rather than plain English and mark up entities such as author, date, and titles rather than syntactic parts of speech. Such entities are generically called "metadata"-data about data-and form an important component of the information present in a bibliography. The aim of our work is to automatically enhance bibliographies with metadata tags, based on a training corpus of annotated bibliography entries.