Efficient Training Data Enrichment and Unknown Token Handling for POS Tagging of Nonstandardized Texts.
Melanie Neunerdt, Michael Reyer, Rudolf Mathar · HilDok – Institutional Repository (Universität Hildesheim) · 2014
In this work we consider the problem of social media text Part-of-Speech tagging as fundamental task for Natural Language Processing. We present improvements to a social media Markov model tagger, by adapting parameter estimation methods for unknown tokens. In addition, we propose to enrich the social media text corpus by a linear combination with a newspaper training corpus. Applying our tagger to a social media text corpus results in accuracies of around 94.8%, which comes close to accuracies for standardized texts.