Exploring the Statistical Derivation of Transformational Rule Sequences for Part-of-Speech Tagging
Lance Ramshaw, Mitchell P. Marcus · 1994
Eric Brill in his recent thesis (1993b) proposed an approach called transformation-based error-driven that can statistically derive linguistic models from corpora, and he has applied the approach in various domains including part-of-speech tagging (Brill, 1992; Brill, 1994) and building phrase structure trees (Brill, 1993a). The method learns a sequence of symbolic rules that characterize important contextual factors and use them to predict a most likely value. The search for such factors only requires counting various sets of events that actually occur in a training corpus, and the method is thus able to survey a larger space of possible contextual factors than could be practically captured by a statistical model that required explicit probability estimates for every possible combination of factors. Brill's results on part-of-speech tagging show that the method can outperform the HMM techniques widely used for that task, while also providing more compact and perspicuo.s models. Decision trees are an established learning technique that is also based on surveying a wide space of possible factors and repeatedly selecting a most significant factor or combination of factors. After briefly describing Brill's approach and noting a fast implementation of it, this paper analyzes it in relation to decision trees. The contrast highlights the kinds of applications to which rule sequence learning is especially suited. We point out how it, ma.ages to largely avoid difficulties with overtraining, and show a way of recording the dependencies bt.tween rules in the learned sequence. The analysis throughout is based on part-of-speech tagging experiments using the tagged Brown Corpus (Francis and K.eera, 1979) and a tagged Septuagint Greek version of the first five books of the Bible (CATSS, 1991). Brill 's Approach