Unify Word-level and Span-level Tasks: NJUNLP’s Participation for the WMT2023 Quality Estimation Shared Task
Xiang Geng, Zhejian Lai, Yu Zhang, Shimin Tao, Hao Yang, Jiajun Chen, Shujian Huang · 2023
We introduce the submissions of the NJUNLP team to the WMT 2023 Quality Estimation (QE) shared task.Our team submitted predictions for the English-German language pair on all two sub-tasks: (i) sentence-and wordlevel quality prediction; and (ii) fine-grained error span detection.This year, we further explore pseudo data methods for QE based on NJUQE framework 1 .We generate pseudo MQM data using parallel data from the WMT translation task.We pre-train the XLMR large model on pseudo QE data, then fine-tune it on real QE data.At both stages, we jointly learn sentence-level scores and word-level tags.Empirically, we conduct experiments to find the key hyper-parameters that improve the performance.Technically, we propose a simple method that covert the word-level outputs to fine-grained error span results.Overall, our models achieved the best results in English-German for both word-level and fine-grained error span detection sub-tasks by a considerable margin.