Internal and external tagsets in part-of-speech tagging

Thorsten Brants · 1997

We present an approach to statistical partof -speech tagging that uses two different tagsets, one for its internal and one for its external representation. The internal tagset is used in the underlying Markov model, while the external tagset constitutes the output of the tagger. The internal tagset can be modified and optimized to increase tagging accuracy (with respect to the external tagset). We evaluate this approach in an experiment and show that it performs significantly better than approaches using only one tagset. 1 Introduction The task of part-of-speech tagging is to assign a unique syntactical category (part-of-speech tag) to each word of an input stream. It is used as a component in parsing, for recognition in message extraction systems, for generating intonation in speech production systems, and many others. Our work focuses on statistical part-of-speech tagging that is based on an underlying n-gram or Markov model. Tags are assigned by maximization of lexica...

Read the paper · More papers on PaperTik