Boosting performance of weak MT engines automatically: using MT output to align segments & build statistical post-editors
Clare R. Voss, Matthew Aguirre, Jeffrey C. Micher, Richard Chang, Jamal Laoudi, Reginald Hobbs · 2008
This paper addresses the practical challenge of improving existing, op- erational translation systems with relatively weak, black-box MT engines when higher quality MT engines are not available and only a limited quantity of online re- sources is available. Recent research results show impressive performance gains in translating between Indo-European languages when mature, existing rule- based MT engines and post-MT editors built automatically with limited amounts of parallel data. We show that this hybrid approach of serially composing or chaining an MT engine and automated post-MT editor---when applied to much weaker lexi- con-based and rule-based MT engines, translating across the more widely divergent languages of Urdu and English, and given limited amounts of document-parallel only training data---will yield statistically significant boosts in translation quality up to the 50K of parallel segments in training the post-editor, but not necessarily be- yond that.