Exploring Representations for Semantic-Rich Part of Speech Tagging

Weidong Qu, Sicong Yue · Advances in computer science research · 2015

Part-of-speech (POS) tagging is the basic and primary analysis step in many natural language processing (NLP) applications.For English, it is often considered a solved problem.There are well established approaches, and the accuracy is around 97% with sufficient domain-specific training data.However, many NLP applications have very different special requirements, and the POS tageset has its own characteristics.These challenges can greatly affect the quality of the part-of-speech tagging process.To address these issues and achieve high POS tagging accuracy, we investigate the representations that can be applied to improve the performance of POS task.Our experiments show that the accuracy of POS tagging degrades significantly when tested with a large semantic and syntactic tagset.In addition, our analysis of experiments suggests that tokens rather than POS tags have more effect on tagging accuracy.Our best results were reached by using the most appropriate representations for POS tagging task.

Read the paper · More papers on PaperTik