Efficient Training Data Enrichment and Unknown Token Handling for POS Tagging of Nonstandardized Texts.

Melanie Neunerdt, Michael Reyer, Rudolf Mathar · HilDok – Institutional Repository (Universität Hildesheim) · 2014

In this work we consider the problem of social media text Part-of-Speech tagging as fundamental task for Natural Language Processing. We present improvements to a social media Markov model tagger, by adapting parameter estimation methods for unknown tokens. In addition, we propose to enrich the social media text corpus by a linear combination with a newspaper training corpus. Applying our tagger to a social media text corpus results in accuracies of around 94.8%, which comes close to accuracies for standardized texts.

Read the paper · More papers on PaperTik