The Impact of Sentence Alignment Errors on Phrase-Based Machine Translation Performance
Cyril Goutte, Marine Jacinthe Carpuat, George Foster · NPARC · 2012
When parallel or comparable corpora are har-vested from the web, there is typically a trade-off between the size and quality of the data. In order to improve quality, corpus collection ef-forts often attempt to fix or remove misaligned sentence pairs. But, at the same time, Statis-tical Machine Translation (SMT) systems are widely assumed to be relatively robust to sen-tence alignment errors. However, there is little empirical evidence to support and character-ize this robustness. This contribution investi-gates the impact of sentence alignment errors on a typical phrase-based SMT system. We confirm that SMT systems are highly tolerant to noise, and that performance only degrades seriously at very high noise levels. Our find-ings suggest that when collecting larger, noisy parallel data for training phrase-based SMT, cleaning up by trying to detect and remove in-correct alignments can actually degrade per-formance. Although fixing errors, when ap-plicable, is a preferable strategy to removal, its benefits only become apparent for fairly high misalignment rates. We provide several expla-nations to support these findings. 1