Recent advances in memory-based part-of-speech tagging

Jakub Zavrel, Walter M. P. Daelemans · Research portal (Tilburg University) · 1999

Memory-based learning algorithms are lazy learners. Examples of a task are stored in memory and processing is largely postponed to the time when new instances of the task need to be solved. This is then done by extrapolating directly from those remembered instances which are most similar to the present ones. Using memory-based learning for Part-of-Speech tagging has a number of advantages over traditional statistical POS taggers: (i) there is no need for an additional smoothing component for sparse data, (ii) even low-frequent or exceptional patterns can contribute to generalization, (iii) the use of a weighted similarity metric allows for an easy integration of different information sources, and (iv) both development time and processing speed are very fast (in the order of hours and thousands of words/sec, respectively). In recent work, we have applied the Memory-Based tagger (MBT) to a number of different languages and corpora (English, Dutch, Czech, Swedish, and Spanish). Furthermor...

Read the paper · More papers on PaperTik