DR-100: Rubric-Based LLM-as-Judge in Machine Translation Via a Simple Meta-Evaluation Framework

Ahrii Kim · 2025

Referred to as LLMs-as-judges, a generative large language model (LLM) has demonstrated considerable efficacy as an evaluator in various tasks, including Machine Translation (MT) evaluation by predicting scores or identifying error types for individual sentences. However, its dependability in practical application has yet to be demonstrated, as there is only an approximated match due to the task's open-ended nature. To address this problem, we introduce a straightforward and novel meta-evaluation framework Overt Assessment and evaluate cutting-edge LLMs such as GEMBA-MQM. We identify their primary obstacles, including certain label biases and the challenge of assessing near-perfect translations. To improve reliability, we investigate more trustworthy and less biased models using multidimensional prompt engineering focused on the skill set. Our findings indicate that the combination of 1 ⃝ span-level error quantification and 2 ⃝ a rubric-style prompt tailored to the characteristics of LLMs has efficiently addressed the majority of the challenges GEMBA faces. Moreover, it has improved the correlation with human judgment from 0.09 (GEMBA) to 0.35, even exceeding that of gold MQM (0.16). Accordingly, we present DR-100, the advanced LLMs-as-judges in MT and an updated version of GEMBA.

Read the paper · More papers on PaperTik