Exploiting rich feature representation for SMT N-best reranking

Yu Ping Tong, Derek F. Wong, Lidia Sam Chao · 2016

N-best reranking in statistical machine translation (SMT) aims to rescore the possible translation hypotheses such that the best hypothesis appears on the top of the list. The current models either use the standard features (a.k.a model features) that are derived from the SMT decoder or the manually crafted features for reranking. However, in practice, the results are disappointing which had only very small effect on the overall performance. In this paper, we investigate using additional semantic and syntactic feature representations to the reranking framework, with the goal to better capture the differences between best hypothesis and others. These representations are the sentence embeddings that are learned using the recursive autoencoder (RAE). The proposed features are extensively evaluated with various reranking algorithms on WMT2015 French-to-English translation data. The empirical results reveal that the proposed reranking model is able to yield an additional improvement of 1.23 BLEU points over the baseline SMT system.

Read the paper · More papers on PaperTik