Artificial intelligence-based automatic evaluation of human translation and interpreting: A systematic review of assessment and validation practices

Chao Han · Research Methods in Applied Linguistics · 2026

• Automatic assessment of translation and interpreting has gained traction. • Machine learning models or machine translation metrics are the mainstay. • Details of human benchmark construction was under-reported. • Score validation relied on correlations with human benchmark scores. • Post-hoc explainability in opaque AI scoring systems was rarely implemented. Human-generated translation and interpreting (T&I) are routinely evaluated in domains such as language education and professional certification. While artificial intelligence (AI) is increasingly used for automatic assessment, little research has examined its application to human T&I. Drawing on rigorous database search and screening, this systematic review attempts to close this gap. Based on a curated corpus of 69 studies, we identify important trends in assessment design, model architecture, and validation practice. The data analysis shows a marked increase in research since 2020, with a dominant focus on English-Chinese T&I, primarily within educational contexts. Most studies employed feature-based machine learning models or repurposed machine translation metrics for scoring, while only a minority explored end-to-end large language models. Benchmark construction was found to be inconsistently reported, with many studies omitting key information about rater qualification, training, reliability, and scoring criteria. Validation practices primarily relied on correlations with human benchmark scores, with limited evidence of convergent validity or cross-condition generalizability. Notably, post-hoc explainability, a crucial step for ensuring transparency in opaque AI systems, was rarely implemented. Overall, this review highlights both progress and persistent challenges in AI-based T&I assessment. While AI holds promise for enhancing assessment efficiency and scalability, methodological limitations and transparency gaps currently constrain its responsible use. We recommend improved reporting standards, multi-pronged validation strategies, development of large annotated benchmark datasets, and greater attention to model interpretability and explainability. These steps are essential for building robust, trustworthy AI systems for automatic T&I assessment.

Read the paper · More papers on PaperTik