Improving POS Tagging Using Machine-Learning Techniques

Lluı́s Màrquez, Horacio Rodríguez, Josep Maria Carmona, Josep Montolio · 1999

In this paper we show how machine learning techniques for constructing and combining several classifiers can be applied to improve the accuracy of an existing English POS tagger (M`arquez and Rodr'iguez, 1997). Additionally, the problem of data sparseness is also addressed by applying a technique of generating convex pseudo--data (Breiman, 1998). Experimental results and a comparison to other state--of--the-- art taggers are reported. Keywords: POS Tagging, Corpus--based modeling, Decision Trees, Ensembles of Classifiers. 1 Introduction The study of general methods to improve the performance in classification tasks, by the combination of different individual classifiers, is a currently very active area of research in supervised learning. In the machine learning (ML) literature this approach is known as ensemble, stacked, or combined classifiers. Given a classification problem, the main goal is to construct several independent classifiers, since it has been proven that when the error...

Read the paper · More papers on PaperTik