Feature Selection for Fluency Ranking

Daniël De Kok, J. Kelleher, B.M. Namee, I. van der Sluis, A. Belz, A. Gatt, A. Koller · 2010

Fluency rankers are used in modern sentence generation systems to pick sentences that are not just grammatical, but also fluent. It has been shown that feature-based models, such as maximum entropy models, work well for this task. Since maximum entropy models allow for incorporation of arbitrary real-valued features, it is often attractive to create very general feature templates, that create a huge number of features. To select the most discriminative features, feature selection can be applied. In this paper we compare three feature selection methods: frequency-based selection, a generalization of maximum entropy feature selection for ranking tasks with realvalued features, and a new selection method based on feature value correlation. We show that the often-used frequency-based selection performs badly compared to maximum entropy feature selection, and that models with a few hundred well-picked features are competitive to models with no feature selection applied. In the experiments described in this paper, we compressed a model of approximately 490.000 features to 1.000 features. is equal to the observed feature value in the training data. In its canonical form, the probability of a certain event (y) occurring in the context (x) is a loglinear combination of features and feature weights, where Z(x) is a normalization over all events in context x (Berger et al., 1996): p(y|x) = 1 Z(x) exp n∑ λifi (1) i=1 The training process estimates optimal feature weights, given the constraints and the principle of maximum entropy. In fluency ranking the input (e.g. a dependency structure) is a context, and a realization of that input is an event within that context. Features can be hand-crafted or generated automatically using very general feature templates. For example, if we apply a template rule that enumerates the rules used to construct a derivation tree to the partial tree in figure 1 the rule(max xp(np)) and rule(np det n) features will be created. 1

Read the paper · More papers on PaperTik