Assessing the effects of translation prompts on the translation quality of GPT-4 Turbo using automated and human evaluation metrics: a case study
Lama Abdullah Aldosari, Nasrin S. Altuwairesh · Perspectives · 2025
The translation industry is transitioning from human-only translation tasks to those increasingly managed by machine translation (MT) systems. While MT has gained attention in legal translation, few studies have focused on using prompt engineering for high-quality translations in this field. This study examines how three existing translation prompts affect the quality of translations of Saudi Laws from Arabic to English generated by the GPT-4 Turbo model. It also explores the impact of incorporating the translation commission framework into new translation prompts. The study involved evaluating translations of fourteen Saudi Laws using both automated (COMET, BLEU, ChrF, TER) and human evaluation metric (MQM). Results show that while existing prompts perform variably, integrating the translation commission into new prompts significantly improves translation quality. The study highlights the importance of collaboration between MT and translation studies researchers to develop better translation prompts for legal texts, contributing to MT, translation studies, computational linguistics, and advancing translation technologies.