Data Selection for Discriminative Training in Statistical Machine Translation

Xingyi Song, Lucia Specia, Trevor Cohn · 2014

The efficacy of discriminative training in Statistical Machine Translation is heavily dependent on the quality of the develop-ment corpus used, and on its similarity to the test set. This paper introduces a novel development corpus selection algo-rithm – the LA selection algorithm. It fo-cuses on the selection of development cor-pora to achieve better translation quality on unseen test data and to make training more stable across different runs, particu-larly when hand-crafted development sets are not available, and for selection from noisy and potentially non-parallel, large scale web crawled data. LA does not re-quire knowledge of the test set, nor the de-coding of the candidate pool before the se-lection. In our experiments, development corpora selected by LA lead to improve-ments of over 2.5 BLEU points when com-pared to random development data selec-tion from the same larger datasets. 1

Read the paper · More papers on PaperTik