CAN WE TRUST MACHINES? A CRITICAL LOOK AT SOME MACHINE TRANSLATION EVALUATION METRICS

Muhammad Zayyanu Zaki, Nazir Ibrahim Abbas · Computer Science & Engineering An International Journal · 2025

The growing interconnection of the globalised world necessitates seamless cross-lingual communication, making Machine Translation (MT) a crucial tool for bridging communication gaps. In this research, the authors have critically evaluated two prominent Machine Translation Evaluation (MTE) metrics: BLEU and METEOR, examining their strengths, weaknesses, and limitations in assessing translation quality, focusing on Hausa-French and Hausa-English translation of some selected proverbs. The authors compared automated metric scores with human judgments of machine-translated text from Google Translate (GT) software. The analysis explores how well BLEU and METEOR capture the nuances of meaning, particularly with culturally bounded expressions of Hausa proverbs, which often have meaning and philosophy. By analysing the performance of the translator’s datasets, they aim to provide a comprehensive overview of the utility of these metrics in Machine Translation (MT) system development research. The authors examined the relationship between automated metrics and human evaluations, identifying where these metrics may be lacking. Their work contributes to a deeper understanding of the challenges of Machine Translation Evaluation (MTE) and suggests potential future directions for creating more robust and reliable evaluation methods. The authors have explored the reasons behind human evaluation of MT quality examining its relationship with automated metrics and its importance in enhancing MT systems. They have analysed the performance of GT as a prominent MT system in translating Hausa proverbs, highlighting the challenges in capturing the language’s cultural and contextual nuances.

Read the paper · More papers on PaperTik