XML encoding for spoken learner (and other) corpora:a modest approach
Andrew Hardie · Kobe University Repository Kernel (Kobe University) · 2014
Since the earliest days of corpus linguistics, markup has been used to represent features of corpus texts other than the actual words of the text. The first systems used were somewhat ad hoc, often based on using (sequences o!l punctuation marks to indicate para-linguistic properties of texts, over time the field has standardized on markup systems based on tags delimited by : first SGML, and more recently the derived XML system. While the official standards for the encoding of corpora with XML - most notably the guidelines of the Text Encoding Initiative - are extremely heavyweight, and therefore most suitable for the development of large-scale reference datasets, I argue that a more modest level of XML can productively be applied within the context of corpora developed by individual researchers or small teams. To apply "Modest XML", it is necessary only to comply with certain fundamental rules of XML , . , and so on}. Finally, it is always possible to extend the XML vocabulary that one uses to support the specific needs of a particular corpus development project.