Literary-Adapted Machine Translation in a Well-Resourced Language Pair

Antonio Toral, Andreas van Cranenburgh, Tia Nutters · 2023

Following recent work on literary-adapted machine translation (MT) systems, this paper investigates whether it is worthwhile building such a system for a reasonably well-resourced language pair, English-to-Dutch, for which generic MT systems (e.g. DeepL) are known to be competitive. Specifically, a system is presented that uses considerably more in-domain training data (novels) than in previous work, as well as an exploration of using longer instances than isolated sentence pairs (i.e. document-level MT). A sizable test set of 31 English-language novels and their published Dutch human translations is evaluated. The evaluation is multidimensional, including automatic MT evaluation metrics, error- and survey-based human evaluation, as well as quantitative automatic analyses, including the novel use of literariness prediction of translations. The results show that, overall, a literary-adapted system that combines sentence- and document-level information performs slightly better than DeepL (4% higher COMET score), with the edge being wider for genre fiction , while the gains over DeepL are smaller or negative for literary fiction . Code, data (public domain subset), and trained systems are available at https://www.w3.org/1999/xlink" xlink:href=" https://github.com/antot/lit-mt-en-nl "> https://github.com/antot/lit-mt-en-nl .

Read the paper · More papers on PaperTik