TMU-HIT’s Submission for the WMT24 Quality Estimation Shared Task: Is GPT-4 a Good Evaluator for Machine Translation?

Ayako Sato, Kyotaro Nakajima, Hwichan Kim, Zhousi Chen, Mamoru Komachi · 2024

In machine translation quality estimation (QE), translation quality is evaluated automatically without the need for reference translations.This paper describes our contribution to the sentence-level subtask of Task 1 at the Ninth Machine Translation Conference (WMT24), which predicts quality scores for neural MT outputs without reference translations.We fine-tune GPT-4o mini, a largescale language model (LLM), with limited data for QE.We report results for the direct assessment (DA) method for four language pairs: English-Gujarati (En-Gu), English-Hindi (En-Hi), English-Tamil (En-Ta), and English-Telugu (En-Te).Experiments under zero-shot, few-shot prompting, and fine-tuning settings revealed significantly low performance in the zero-shot, while fine-tuning achieved accuracy comparable to last year's best scores.Our system demonstrated the effectiveness of this approach in low-resource language QE, securing 1st place in both En-Gu and En-Hi, and 4th place in En-Ta and En-Te.The code used in our experiments is available at the following URL 1 .

Read the paper · More papers on PaperTik