Learning a Phrase-based Translation Model from Mon- olingual Data with Application to Domain Adaptation

Jiajun Zhang, Chengqing Zong · 2015

Currently, almost all of the statistical ma-chine translation (SMT) models are trained with the parallel corpora in some specific domains. However, when it comes to a lan-guage pair or a different domain without any bilingual resources, the traditional SMT loses its power. Recently, some research works study the unsupervised SMT for in-ducing a simple word-based translation model from the monolingual corpora. It successfully bypasses the constraint of bitext for SMT and obtains a relatively promising result. In this paper, we take a step forward and propose a simple but effec-tive method to induce a phrase-based model from the monolingual corpora given an au-tomatically-induced translation lexicon or a manually-edited translation dictionary. We apply our method for the domain adaptation task and the extensive experiments show that our proposed method can substantially improve the translation quality. 1

Read the paper · More papers on PaperTik