Monolingual Marginal Matching for Translation Model Adaptation

Ann Irvine, Chris Quirk, Hal Daumé · 2013

When using a machine translation (MT) model trained on OLD-domain parallel data to translate NEW-domain text, one major challenge is the large number of out-of-vocabulary (OOV) and new-translation-sense words.We present a method to identify new translations of both known and unknown source language words that uses NEW-domain comparable document pairs.Starting with a joint distribution of source-target word pairs derived from the OLD-domain parallel corpus, our method recovers a new joint distribution that matches the marginal distributions of the NEW-domain comparable document pairs, while minimizing the divergence from the OLD-domain distribution.Adding learned translations to our French-English MT model results in gains of about 2 BLEU points over strong baselines.

Read the paper · More papers on PaperTik