Large Language Models Evaluate Machine Translation via Polishing
Yiheng Wang · 2023
Quality of machine translation is usually evaluated by automatic reference-based metrics. However, these conventional metrics have limited correlation with human evaluation and require references for comparison. Human evaluation is considered the most reliable, but it is still plagued by high costs and uncertainties among annotators. With the emergence of large language models (LLMs) and their impressive capabilities, some studies have utilized LLMs for the evaluation of natural language generation. In this paper, we propose EvLP (Evaluation via LLMs Polishing), a reference-free machine translation evaluation method inspired by post-editing that leverages LLMs. We select gpt-3.5-turbo to act as annotators to polish the translated text. Our experiments compare EvLP with other LLM-based evaluation methods and SEG_LM, a representative reference-free metric utilizing cross-lingual pre-trained language models to evaluate machine translation. We select a dataset with two kinds of human evaluation, direct assessment and post-editing. The results of experiments demonstrate that EvLP closely aligns with human evaluation compared to other LLM-based metrics, outperforming SEG_LM on direct assessment and exhibiting a significant advantage in post-editing. We also discuss the bias that LLMs hold towards LLM-generated texts.