Improving Machine Translation Quality Prediction with Syntactic Tree Kernels
Christian Hardmeier · 2011
We investigate the problem of predicting the quality of a given Machine Translation (MT) output segment as a binary classifi-cation task. In a study with four different data sets in two text genres and two lan-guage pairs, we show that the performance of a Support Vector Machine (SVM) clas-sifier can be improved by extending the feature set with implicitly defined syn-tactic features in the form of tree ker-nels over syntactic parse trees. Moreover, we demonstrate that syntax tree kernels achieve surprisingly high performance lev-els even without additional features, which makes them suitable as a low-effort initial building block for an MT quality estima-tion system. 1