Translation Model Adaptation Using Genre-Revealing Text Features

Marlies van der Wees, Arianna Bisazza, Christof Monz · 2015

Research in domain adaptation for statistical machine translation (SMT) has resulted in various approaches that adapt system components to specific translation tasks.The concept of a domain, however, is not precisely defined, and most approaches rely on provenance information or manual subcorpus labels, while genre differences have not been addressed explicitly.Motivated by the large translation quality gap that is commonly observed between different genres in a test corpus, we explore the use of document-level genrerevealing text features for the task of translation model adaptation.Results show that automatic indicators of genre can replace manual subcorpus labels, yielding significant improvements across two test sets of up to 0.9 BLEU.In addition, we find that our genre-adapted translation models encourage document-level translation consistency.

Read the paper · More papers on PaperTik