Using Comparable Corpora to Adapt MT Models to New Domains
Ann Irvine, Chris Callison-Burch · 2014
In previous work we showed that when using an SMT model trained on old-domain data to translate text in a new-domain, most errors are due to unseen source words, unseen target translations, and inaccurate translation model scores (Irvine et al., 2013a).In this work, we target errors due to inaccurate translation model scores using new-domain comparable corpora, which we mine from Wikipedia.We assume that we have access to a large olddomain parallel training corpus but only enough new-domain parallel data to tune model parameters and do evaluation.We use the new-domain comparable corpora to estimate additional feature scores over the phrase pairs in our baseline models.Augmenting models with the new features improves the quality of machine translations in the medical and science domains by up to 1.3 BLEU points over very strong baselines trained on the 150 million word Canadian Hansard dataset.