Improved Statistical Machine Translation with Source Language Paraphrase

Su Che · Beijing Daxue Xuebao. Zirankexueban · 2015

The performance of statistical machine translation(SMT) suffers from the insufficiency of parallel corpus. To solve the problem, the authors propose a paraphrase based SMT framework with three solutions: 1) acquiring paraphrase knowledge based on a third language; 2) expressing multiple paraphrases of input sentence in a lattice and modifying decoder to be able to process it; 3) integrating paraphrase knowledge as features into loglinear model. In this way, not only more expressions in source language can be covered, but also more expressions in target language can be generated as candidate translations. To verify proposed method, experiments are conducted on three training data sets with different sizes, and evaluate the improvement of the performance of SMT system contributed by paraphrasing. Experimental results show that the translation performance is improved significantly(BLEU+1.4%) when the parallel corpus is small(10 K), and a good performance(BLEU+0.32%) is also achieved when parallel corpus is large enough(1 M).

Read the paper · More papers on PaperTik