Tuning SMT with a Large Number of Features via Online Feature Grouping

Lemao Liu, Tiejun Zhao, Taro Watanabe, Eiichiro Sumita · 2013

In this paper, we consider the tuning of sta-tistical machine translation (SMT) mod-els employing a large number of features. We argue that existing tuning methods for these models suffer serious sparsity prob-lems, in which features appearing in the tuning data may not appear in the test-ing data and thus those features may be over tuned in the tuning data. As a result, we face an over-fitting problem, which limits the generalization abilities of the learned models. Based on our analysis, we propose a novel method based on feature grouping via OSCAR to overcome these pitfalls. Our feature grouping is imple-mented within an online learning frame-work and thus it is efficient for a large scale (both for features and examples) of learning in our scenario. Experiment re-sults on IWSLT translation tasks show that the proposed method significantly outper-forms the state of the art tuning methods. 1

Read the paper · More papers on PaperTik