Linguistic Indicators for Quality Estimation of Machine Translations

Mariano Felice · 2012

This work presents a study of linguistically-informed features for the automatic quality estimation of machine translations. In particular, we address the problem of estimating quality when no reference translations are available, as this is the most common case in real world situations. Unlike previous attempts that make use of internal information from translation systems or rely on purely shallow aspects, our approach uses features derived from the source and target text as well as additional linguistic resources, such as parsers and monolingual corpora. We built several models using a supervised regression algorithm and different combinations of features, contrasting purely shallow, linguistic and hybrid sets. Evaluation of our linguistically-enriched models yields mixed results. On the one hand, all our hybrid sets beat a shallow baseline in terms of Mean Average Error but on the other hand, purely linguistic feature sets are unable to outperform shallow features. However, a detailed analysis of individual feature performance and optimal sets obtained from feature selection reveals that shallow and linguistic features are in fact complementary and must be carefully combined to achieve optimal results. In effect, we demonstrate that the best performing models are actually based on hybrid sets having a significant proportion of linguistic features. Furthermore, we show that linguistic information can produce consistently better quality estimates for specific score intervals. Finally, we analyse many factors that may have an impact on the performance of linguistic features and suggest new directions to mitigate them in the future.

Read the paper · More papers on PaperTik