Quality Estimation for Extending Good Translations

Lucia Specia, Kashif Ur Rehman Shah · 2014

We present experiments using quality estimation models to improve the performance of statistical machine translation (SMT) systems by supplementing their training corpora or by building sentence-specific SMT models for instances predicted as having potential for improve- ment by the ITERPE model. The experiments with quality-informed active learning strategy select, among alternative machine translations, those which are: (i) predicted to have high qual- ity, and thus can be added to the machine translation system training set; (ii) predicted to have low quality, and thus need to be corrected/translated by humans, with the human corrections added to the machine translation system training set. Improvement is measured by the increase in the performance of the overall machine translation systems on held-out datasets, where perfor- mance is measured by automatic evaluation metrics comparing the scores of the original machine translation system against the score of the improved machine translation system after additional material is used. The experiments with ITERPE consist in automatically grouping translation in- stances into different quality bands, for instance for re-translation or for post-editing (Bicici and Specia, 2014). This method can be helpful in automatic identification of quality barriers in MT to achieve high quality machine translation.

Read the paper · More papers on PaperTik