Automatically acquiring a language model for POS tagging using decision trees

Lluı́s Màrquez, Horacio Rodríguez · Amsterdam studies in the theory and history of linguistic science. Series 4, Current issues in linguistic theory · 2000

We present an algorithm that automatically acquires a statistically--based language model for POS tagging, using statistical decision trees. The learning algorithm deals with more complex contextual information than simple collections of n--grams and it is able to use information of different nature. The acquired models are independent enough to be easily incorporated, as a statistical core of constraints/rules, in any flexible tagger. They are also complete enough to be directly used as sets of POS disambiguation rules. We have implemented a simple and fast tagger that has been tested and evaluated on the WSJ corpus with a remarkable accuracy. Comparative results are reported. 1 Introduction In NLP, it is necessary to model the language in a representation suitable for the task to be performed. The language models more commonly used are based on two main approaches: first, the linguistic approach, in which the model is written by a linguist, generally in the form of rules or constrai...

Read the paper · More papers on PaperTik