A cross-language methodology for corpus part-of-speech tag-set development
Eric Atwell · White Rose Research Online (University of Leeds, The University of Sheffield, University of York) · 2007
This paper examines criteria used in development of Corpus Part-of-Speech tag sets used when PoS-tagging a corpus, that is, enriching a corpus by adding a part-ofspeech category label to each word. This requires a tag-set, a list of grammatical category labels; a tagging scheme, practical definitions of each tag or label, showing words and contexts where each tag applies; and a tagger, a program for assigning a tag to each word in the corpus, implementing the tag-set and tagging-scheme in a tagassignment algorithm. We start by reviewing tag-sets developed for English corpora, since English was the first language studied by corpus linguists. Traditional English grammars generally provide 8 basic parts of speech, derived from Latin grammar. However, most tag-set developers wanted to capture finer grammatical distinctions, leading to larger tag-sets. Figure 1 illustrates a range of rival English PoS-tag-sets applied to a short example sentence; even with this simple sentence, it is easy to see some significant similarities and differences between these rival tag-sets for English. The pioneering Corpus Linguists who collected the first large-scale English language corpora all thought that their corpora could be more useful research resources if the source text samples were enriched with linguistic analyses. These pioneering English corpus linguistics projects included projects to collect the Brown corpus, the