Log-Linear Reformulation of the Noisy Channel Model for Document-Level Neural Machine Translation
Sébastien Jean, Kyunghyun Cho · 2020
We seek to maximally use various data sources, such as parallel and monolingual data, to build an effective and efficient documentlevel translation system.In particular, we start by considering a noisy channel approach (Yu et al., 2020) that combines a target-to-source translation model and a language model.By applying Bayes' rule strategically, we reformulate this approach as a log-linear combination of translation, sentence-level and documentlevel language model probabilities.In addition to using static coefficients for each term, this formulation alternatively allows for the learning of dynamic per-token weights to more finely control the impact of the language models.Using both static or dynamic coefficients leads to improvements over a context-agnostic baseline and a context-aware concatenation model.