Markup Systems

Gerald Nelson · 1996

Abstract Markup is the first level of annotation applied to the component corpora in ICE. It may be divided into two distinct types: textual markup, which is added to the texts themselves, and bibliographical and biographical markup, which is stored extern ally in the form of a file header for each text. The system for textual markup is based on a proposal by Rosta (1990) and is fully described in two manuals, one each for spoken and written texts (Nelson, 1991a, 1991b). The system for encoding bibliographical and biographical information is described in Nelson (1991c). In this paper I will discuss both markup types in tum, giving examples from the British ICE corpus (ICE-GB). Finally, I will discuss some of the ways in which markup is used in text retrieval.

Read the paper · More papers on PaperTik