Analysing Translation Artifacts: A Comparative Study of LLMs, NMTs, and Human Translations

Fedor Sizov, Cristina España-Bonet, Josef van Genabith, Roy Xie, Koel Dutta Chowdhury · 2024

Translated texts exhibit a range of characteristics that make them appear distinct from texts originally written in the same target language.With the rise of Large Language Models (LLMs), which are designed for a wide range of language generation and understanding tasks, there has been significant interest in their application to Machine Translation.While several studies have focused on improving translation quality through fine-tuning or few-shot prompting techniques, there has been limited exploration of how LLM-generated translations qualitatively differ from those produced by Neural Machine Translation (NMT) models, and human translations.Our study employs explainability methods such as Leave-One-Out (LOO) and Integrated Gradients (IG) to analyze the lexical features distinguishing human translations from those produced by LLMs and NMT systems.Specifically, we apply a twostage approach: first, classifying texts based on their origin -whether they are original or translations-and second, extracting significant lexical features (highly attributed input words) using post-hoc interpretability methods.Our analysis shows that different methods of feature extraction vary in their effectiveness, with LOO being generally better at pinpointing critical input words and IG capturing a broader range of important words.Finally, our results show that while LLMs and NMT systems can produce translations of a good quality, they still differ from texts originally written by native speakers.We find that while some LLMs more closely resemble human translations, traditional NMT systems show distinct differences, particularly in their use of linguistic features.1

Read the paper · More papers on PaperTik