Corpus encoding and annotation

Federico Zanettin · 2014

Annotation is but one part of corpus construction, which can contribute to the creation of stable, flexible and accessible corpus resources for translation studies. The annotation framework, however, allows for the addition of further layers of annotation, which can be carried out using resources. This chapter examines how the translation-driven corpora, either monoor multilingual, can be enriched with documentary, structural and linguistic annotation. The annotation of a corpus using Extensible Markup Language (XML), Text Encoding Initiative (TEI) and XML Corpus Encoding Standard (XCES) certainly requires an understanding of how annotation schemes work. The chapter analyses how to annotate translation-driven corpora within the TEI annotation framework, and proposes some practical tasks which take advantage of a freely distributed and compatible corpus analysis program, namely XML Aware Information Retrieval Architecture (XAIRA). It explains how the TEI XML annotation standard, which provides guidelines for the annotation of large language corpora, may offer a suitable annotation scheme for translation-driven corpora.

Read the paper · More papers on PaperTik