Feasibility of automatic evaluation metrics as quantitative assessment tools for Korean translation texts: focused on SacreBLEU

Jisoo Choi · Language and Information · 2023

In this study, I conducted an examination of the effectiveness of automated evaluation metrics, primarily designed for machine translation, in the realm of human translation assessment. The research involved a thorough comparison of these automated metrics, such as BLEU and METEOR, with conventional manual evaluation techniques. Central to this investigation was the application of SacreBLEU, a refined metric with standardized preprocessing, to evaluate human-generated texts. The findings revealed a notable correlation, exceeding 0.8, between the outcomes of automated evaluations and human judgments. This result highlights the potential of automated metrics to enhance the accuracy and efficiency of human translation assessment, and it is easy to use, further emphasizing its practical utility. The study thus advocates for the integration of these metrics into academic and professional translation evaluation practices, offering a more streamlined, consistent, and objective approach to translation quality assessment.

Read the paper · More papers on PaperTik