Marking Words with Part-of-Speech (POS) Tags within Text Boundary of a Corpus: the Problems, the Process and the Outcomes
Translation Today · 2015
A natural language text stored in a corpus database in electronic version can be tagged at the part-of-speech (POS) level manually or automatically.In both cases, it has to be done carefully starting with the lowest level of hierarchy of tagset meticulously devised for a language or a language group.Once the lower level tag is selected and assigned to words, the higher level tags will be automatically identified and assigned.Although tagging of words may be done with a focus on the part-of-speech of words used in a piece of text, the long term goals should also be envisaged for developing a generic scheme that may be useful for incorporating various kinds of linguistic information easily at the later stages of text annotation.This paper argues for taking a judicious decision for tagging words with different types of information within a text following the universally accepted principles, maxims and rules adopted for part-of-speech tagging.It describes the strategies, rules and methods adopted for manual tagging of a Bengali written text corpus at the part-of-speech level following the guidelines and methods proposed in the Bureau of Indian Standard (BIS) suitable for the language.