Quality Scoring of Source Words in Neural Translation Models
Priyesh Jain, Sunita Sarawagi, Tushar Tomar · 2022
Word-level quality scores on input source sentences can provide useful feedback to an enduser when translating into an unfamiliar target language.Recent approaches either require training custom models on synthetic data or repeatedly invoking the translation model.We propose a simple approach based on comparing probabilities from two language models.The basic premise of our method is to reason how well each source word is explained by the generated translation as against the preceding source language words.Our approach provides between 2.2 and 27.1 higher F1 score and is significantly faster than state of the art methods on three language pairs.Also, our method does not require training any new model.We release a public dataset on word omissions and mistranslations on a new language pair. 1