Predictive modeling of error categories in English-Slovak machine translation using automatic evaluation metrics
Daša Munková, Lucia Benková, Michal Munk, Ľubomír Benko, Petr Hájek · Machine Learning with Applications · 2025
• Automated methodology for identifying MT error types. • The methodology offers a scalable framework for MT quality evaluation across languages. • The modeling techniques predict the likelihood of specific error types. This paper presents a language-specific adaptation for automatic identification of machine translation (MT) errors using a comprehensive set of open-source evaluation metrics. The approach focuses on the English–Slovak translation direction, addressing challenges posed by Slovak’s highly inflectional and low-resource nature. Predictive models were developed for five key error categories (Predication, Modal and communication sentence framework, Syntactic-semantic correlativeness, Compound/complex sentences, and Lexical semantics) by employing forward stepwise regression and validated through bootstrapping techniques. The models estimate the probability of error occurrence in MT segments, demonstrating stable and comparable performance across training and test datasets, as measured by Somers’ D. While human expert evaluation remains essential for verifying flagged segments, the proposed approach significantly reduces evaluator workload by prioritizing likely error-containing segments. This methodology offers a scalable and adaptable framework for MT quality assessment across languages and text styles, with potential to improve automated translation evaluation and post-editing processes.