This Phrase-Based SMT System is Out of Order: Generalised Word Reordering in Machine Translation

Simon Zwarts, Mark Dras · 2006

Many natural language processes have some degree of preprocessing of data: tokenisation, stemming and so on. In the domain of Statisti-cal Machine Translation it has been shown that word reordering as a preprocessing step can help the translation process. Recently, hand-written rules for reordering in German–English translation have shown good results, but this is clearly a labour-intensive and language pair-specific approach. Two possible sources of the observed improvement are that (1) the reordering explicitly matches the syntax of the source language more closely to that of the target language, or that (2) it fits the data bet-ter to the mechanisms of phrasal SMT; but it is not clear which. In this paper, we apply a gen-eral principle based on dependency distance min-imisation to produce reorderings. Our language-independent approach achieves half of the im-provement of a reimplementation of the hand-crafted approach, and suggests that reason (2) is a possible explanation for why that reordering ap-proach works. Help you I can, yes. Jedi Master Yoda 1

Read the paper · More papers on PaperTik