Selection of SMT Training Data Based on Sentence Pair Quality and Coverage
Shujie Yao · Zhongwen xinxi xuebao · 2011
In Statistical Machine Translation,effective selection of training data can generally reduce the burden of system training and decoding.To addressing this issue,,we propose a framework to select a small portion from the whole training data set for SMT by considering both coverage and sentence pair quality.Experimental results on CWMT2008 Chinese-to-English MT task show that our framework is effective to select a subset from the large training data set.Even trained on the 20% data selected by our framework,the SMT system can achieve comparable performance with the baseline system trained on all the data).