Data enhancement and selection strategies for the word-level Quality Estimation
Varvara Logacheva, Chris Hokamp, Lucia Specia · 2015
This paper describes the DCU-SHEFF word-level Quality Estimation (QE) system submitted to the QE shared task at WMT15.Starting from a baseline set of features and a CRF algorithm to learn a sequence tagging model, we propose improvements in two ways: (i) by filtering out the training sentences containing too few errors, and (ii) by adding incomplete sequences to the training data to enrich the model with new information.We also experiment with considering the task as a classification problem, and report results using a subset of the features with Random Forest classifiers.