Part-of-Speech Tagging of Dutch with MBT, a Memory-Based Tagger Generator
Walter M. P. Daelemans, Jakub Zavrel · 2007
We present a part of speech tagger (morphosyntactic disambiguator) for Dutch, constructed by means of the Memory-Based Tagger generation method. In this approach, inductive learning methods are used to derive a tagger, lexicon and unknown word category guesser fully automatically from a tagged example corpus. Advantages of the approach are (i) fast tagger development time without linguistic engineering, (ii) accuracy better than or comparable to state of the art statistical and rule-based approaches, (iii) fast tagging speed, and (iv) reliable unknown word category guessing without the overhead of morphological analysis. 1 Introduction A Part-of-Speech tagger annotates the words in a text with their morphosyntactic categories. A good tagger is instrumental in a large number of information technology solutions. It can produce a shallow, but accurate linguistic analysis of texts, and therefore features as a central component in many text processing and language engineering applications ...