Context is Ubiquitous, but Rarely Changes Judgments: Revisiting Document-Level MT Evaluation
Ahrii Kim · 2025
As sentence-level performance in modern Machine Translation (MT) models reaches a plateau where differences are minimal, there is a growing need for robust document-level evaluation methods. We present a reproducible human evaluation protocol that is structured upon the FALCON framework (Kim, 2025) encompassing pragmatic features. With professional translators as annotators, we investigate the sources of low inter-annotator agreement and identify the primary contributing factors. To address these challenges and align with human values, we propose a comprehensive annotation-rating methodology referred to as H-FALCON. Our experiment shows that, while perfect annotator consensus remains elusive, the proposed scoring scheme achieves equal or higher correlations with traditional sentencelevel metrics. Linear regression analysis further reveals that contextual information is inherent in all sentences-contrary to the belief that only a subset requires it-and that previous estimates such as "n % of sentences require context" stem from flawed calculations. Context contributes approximately 10% to the variance of the holistic score in our evaluation, highlighting its universal yet limited influence on the MT evaluation. Codes will be released.