Reordered Search, and Tuple Unfolding for Ngram-based SMT.

Josep Crego, José Bernardo Mariño Acebal, Adrià de Gispert · 2005

In Statistical Machine Translation, the use of reordering for certain language pairs can pro-duce a significant improvement on translation accuracy. However, the search problem is shown to be NP-hard when arbitrary reorderings are allowed. This paper addresses the question of reordering for an Ngram-based SMT approach following two complementary strategies, namely reordered search and tuple unfolding. These strategies interact to improve translation qual-ity in a Chinese to English task. On the one hand, we allow for an Ngram-based decoder (MARIE) to perform a reordered search over the source sentence, while combin-ing a translation tuples Ngram model, a tar-get language model, a word penalty and a word distance model. Interestingly, even though the translation units are learnt sequentially, its re-ordered search produces an improved transla-tion. On the other hand, we allow for a modifica-tion of the translation units that unfolds the tuples, so that shorter units are learnt from a new parallel corpus, where the source sen-tences are reordered according to the target lan-guage. This tuple unfolding technique reduces data sparseness and, when combined with the reordered search, further boosts translation per-formance. Translation accuracy and efficency results are reported for the IWSLT 2004 Chinese to English task. 1

Read the paper · More papers on PaperTik