Part-of-Speech Tagging by Means of ILP and Active Learning

Miloslav Nepil, Lubomír Popelínský · 2001

Morphological tagging in inflectional languages like Czech is quite a difficult problem that can be hardly solved without employing (semi-)automatic taggers. Several morphological taggers have been developed (or proposed) for Czech, based on either statistical methods [5, 8] or machine learning techniques [9, 10]. Here we focus on part-of-speech (POS) tagging. For instance, in the two sentences in Fig. 1, the word form own can be either adjective or verb. Our goal (oni) !vlastn#? auto. (They !own? a car.) !Zni#ili jejich !vlastn#? auto. (They destroyed their !own? car.) Fig. 1. Example of ambiguous words is to nd the correct POS for a given word form if we know the words in its context. To build a tagger one usually needs a representative set of texts. The principal difficulty for Czech lies in the fact that annotated corpora (i.e. unambiguously tagged ones) are too small. Compare 164 000 st...

Read the paper · More papers on PaperTik