COMET-poly: Machine Translation Metric Grounded in Other Candidates

Maike Züfle, Vilém Zouhar, Tu Anh Dinh, Felipe Maia Polo, Jan Niehues, Mrinmaya Sachan · 2025

Automated metrics for machine translation attempt to replicate human judgment.Unlike humans, who often assess a translation in the context of multiple alternatives, these metrics typically consider only the source sentence and a single translation.This discrepancy in the evaluation setup may negatively impact the performance of automated metrics.We propose two automated metrics that incorporate additional information beyond the single translation.COMET poly-cand uses alternative translations of the same source sentence to compare and contrast with the translation at hand, thereby providing a more informed assessment of its quality.COMET poly-ic , inspired by retrieval-based in-context learning, takes in translations of similar source texts along with their human-labeled quality scores to guide the evaluation.We find that including a single additional translation in COMET poly-cand improves the segment-level metric performance (0.079→0.118 τ b ), with further gains when more translations are added.Incorporating retrieved examples in COMET poly-ic yields similar improvements (0.079→0.116 τ b ).We release our models publicly.1

Read the paper · More papers on PaperTik