Document-level machine translation evaluation project: methodology,effort and inter-annotator agreement

Sheila Castilho · 2020

Recently, document-level (doc-level) human evaluation of machine translation (MT) has raised interest in the community after a few attempts have disproved claims of “human parity” (Toral et al., 2018; Laubli et al., 2018). However, lit- ¨ tle is still known about best practices regarding doc-level human evaluation. This project aims to identify methodologies to better cope with i) the current state-of-theart (SOTA) human metrics, ii) a possible complexity when assigning a single score to a text consisted of ‘good’ and ‘bad’ sentences, iii) a possible tiredness bias in doc-level set-ups, and iv) the difference in inter-annotator agreement (IAA) between sentence and doc-level set-ups.

Read the paper · More papers on PaperTik