FUN-NRC: Paraphrase-augmented Phrase-based SMT Systems for NTCIR-10 PatentMT
Atsushi Fujita, Marine Jacinthe Carpuat · 2013
This paper describes FUN-NRC group’s machine translation sys-tems that participated in the NTCIR-10 PatentMT task. The central motivation of this participation was to clarify the potential of auto-matically compiled collections of sub-sentential paraphrases. Our systems were built using our baseline phrase-based SMT system by augmenting its phrase table with novel translation pairs gener-ated by combining paraphrases with translation pairs learned di-rectly from the training bilingual data. We investigated two meth-ods for phrase table augmentation: source-side augmentation and target-side augmentation. Among the systems we submitted, the two that worked best were (a) the one that paraphrased only unseen phrases into translatable phrases at the source side and (b) the one that paraphrased target phrases only into phrases that were seen in the original phrase table. Both these systems were trained on not only bilingual, but also monolingual data. The other two systems were trained using only bilingual data. This paper also reports on our follow-up experiments focusing on the relationship between re-ordering restriction and system performance.