Machine Translation Metrics Are Better in Evaluating Linguistic Errors on LLMs than on Encoder-Decoder Systems

Eleftherios Avramidis, Shushen Manakhimova, Vivien Macketanz, Sebastian Möller · 2024

This year's MT metrics challenge set submission by DFKI expands previous years' linguistically motivated challenge set.It includes 137,000 items extracted from 100 MT systems for the two language directions (en!de, en!ru), covering more than 100 linguistically motivated phenomena organized in 14 linguistic categories.The metrics with the statistically significant best performance with regard to our linguistically motivated analysis are METRICX-24-HYBRID and METRICX-24 for en!de and METRICX-24 for en!ru, whereas METAMET-RICS and XCOMET are in the next ranking positions in both language pairs.Metrics are more accurate in detecting linguistic errors among LLM translations than in translations based on the encoder-decoder NMT architecture.Some of the most difficult phenomena for the metrics to score are the transitive past progressive, the multiple connectors, and the ditransitive simple future I for en!de and the pseudogapping, the contact clause and the cleft sentences for en!ru.Despite its overall low performance, the LLM-based metric GEMBA performs best in scoring German negation errors.

Read the paper · More papers on PaperTik