Automatic Tune Set Generation for Machine Translation with Limited Indomain Data
Jinying Chen, Jacob Devlin, Huaigu Cao, Rohit Prasad, Prem Natarajan · 2012
Many effective adaptation techniques for statistical machine translation crucially rely on in-domain development sets to learn model parameters. In this paper we present a novel method that automatically generates the matching tune set for Arabic-to-English MT with limited indomain data 1 . This technique improves our MT system over two baselines (tuned on data from the same domain but different genres) by 1.2 and 3.5 BLEU points using significantly less tuning data (1/6 and 1/2 of the baselines). Lexical and morphological features contribute to the success of our method in different ways. Generating tune sets using length distribution also improves the system significantly. Finally, our method obtains competitive results in experiments where ingenre tune sets are available.