A Comparison of Criteria for Maximum Entropy/ Minimum Divergence Feature Selection

Adam L. Berger, Harry Printz · 1998

In this paper we study the gain, a naturally-arising statistic from the theory of memd modeling [2], as a figure of merit for selecting features for an memd language model. We compare the gain with two popular alternatives---empirical activation and mutual information---and argue that the gain is the preferred statistic, on the grounds that it directly measures a feature 's contribution to improving upon the base model. Introduction Maximum entropy / minimum divergence (memd) modeling is a powerful technique for building statistical models of linguistic phenomena. It has been applied to problems as diverse as machine translation [2], parsing [10], word morphology [5] and language modeling [6, 11, 3, 9]. The heart of the method is to choose a collection of informative features, each encoding some linguistically significant event, and then to incorporate these features into a family of conditional models. A fundamental issue in applying this technique is the criterion used to select f...

Read the paper · More papers on PaperTik