An Empirical Comparison of Features and Tuning for Phrase-based Machine Translation
Spence Green, Daniel M. Cer, Christopher D. Manning · 2014
Scalable discriminative training methods are now broadly available for estimating phrase-based, feature-rich translation models.However, the sparse feature sets typically appearing in research evaluations are less attractive than standard dense features such as language and translation model probabilities: they often overfit, do not generalize, or require complex and slow feature extractors.This paper introduces extended features, which are more specific than dense features yet more general than lexicalized sparse features.Large-scale experiments show that extended features yield robust BLEU gains for both Arabic-English (+1.05) and Chinese-English (+0.67) relative to a strong feature-rich baseline.We also specialize the feature set to specific data domains, identify an objective function that is less prone to overfitting, and release fast, scalable, and language-independent tools for implementing the features.Lexicalized rule indicator (Liang et al., 2006a) Some rules occur frequently enough that we can learn rule-specific weights that augment the dense translation model features.For example, our model learns the following rule indicator features and weights: ⇒ reasons -0.022 ⇒ reasons for 0.002 ⇒ the reasons for 0.016