Decision Tree Ensemble for Parts-of-Speech Tagging of Resource-poor Languages
Vamshi K. G. Reddy, Pratibha Rani, Vikram Pudi, Dipti Misra Sharma · 2018
Ensemble POS taggers are a good choice to integrate and leverage benefits of various types of POS taggers. This can help the large number (6500+) of resource-poor languages which do not have much annotated training data by providing ways to integrate semi-supervised/unsupervised taggers with supervised taggers. In this paper we present our experiments of developing ensemble POS taggers using a decision tree. We integrate a semi-supervised data mining approach that uses context based lists (CBLs) for POS tagging with supervised (1) Support Vector Machine based POS tagger, called SVMTool and (2) Conditional Random Field based POS tagger. The results are enhanced semi-supervised ensemble POS taggers which outperform the base methods. In these POS taggers, we use a decision tree to decide when to rely on the output of supervised tagger, and when to rely on the semi-supervised CBL method. The CBL based tagger uses rich contextual information which helps in tagging both existing and unseen words and uses no domain knowledge while supervised taggers give good performance for words present in the training model and can include domain based features. Hence, these algorithms have complementary strengths and in our ensemble we are able to combine these strengths. Enhanced performance of our new POS taggers over the base methods suggests that integrating these methods combines the qualities of these in the new tagger which enhances the performance. Therefore, these new semi-supervised ensemble taggers are more suitable for resource-poor languages.