Data for Machine Translation

Joss Moorkens, Andy Way, Séamus Lankford · 2024

This chapter demonstrates the critical role played by data in machine learning (ML) approaches to machine translation (MT). We examine how data can be sourced, and describe the differences between authentic human-produced data and synthetic data generated by a machine. A number of data-related issues are covered, including alignment, formatting, bias, and ownership. We show very clearly that the playing field is not a level one -- for most languages and language-pairs, insufficient data exists for good-quality corpus-driven translation models to be built; for spoken and sign languages, the situation is even worse, so any talk of MT being a ‘solved problem’ is hugely premature. Where data does exist in sufficient amounts, we describe pre- and post-processing routines that are required prior to the training of ML models.

Read the paper · More papers on PaperTik