Approach to Selecting Best Development Set for Phrase-Based Statistical Machine Translation
Peng Liu, Zhou Yu, Chengqing Zong · Institutional Repositories DataBase (IRDB) · 2015
Abstract. In phrase-based statistical machine translation system, the parameters of model are usually obtained from minimum error rate training (MERT) on the development set. So the development set has a great influence on the performance of the translation system. Generally, more development set will achieve more effective and robust parameters but consume much more time on MERT process. In this paper, we propose two methods to select sentences from the large development set, based on the phrase and the sentence structure respectively. The experimental results show that our methods can get better translation performance than the baseline system on the compact development set by using a state-of-the-art SMT system.