A joint translation model with integrated reordering

Nadir Durrani · OPUS Publication Server of the University of Stuttgart (University of Stuttgart) · 2012

This dissertation aims at combining the benefits and to remedy the flaws of the two popular frameworks in statistical machine translation, namely Phrase-based MT and N-gram-based MT. Phrase-based MT advanced the state-of-the art towards translating phrases than words. By memorizing phrases, phrasal MT, is able to learn local reorderings, and handling of other local dependencies such as insertions, deletions etc. Inter-phrasal reorderings are handled through the lexicalized reordering model, which remains the state-of-the-art model for reordering in phrase-based SMT till date. However, phrase-based MT has some drawbacks: • Dependencies across phrases are not directly represented in the translation model • Discontinuous phrases cannot be represented and used • The reordering model is not designed to handle long range reorderings • Search and modeling problems require the use of a hard reordering limit • The presence of many different equivalent segmentations increases the search space • Source word deletion and target word insertion outside phrases is not allowed during decoding N-gram-based MT exists as an alternate to the more commonly used Phrase-based MT. Unlike Phrasal MT, N-gram-based MT uses minimal translation units called as tuples. Using minimal translation units, enables N-gram systems to avoid the spurious phrasal segmentation problem in the phrase-based MT. However, it also gives up the ability to memorize dependencies such as short reorderings that are local to the phrases. Reordering in N-gram MT is carried out by source linearization and POS-based rewrite rules. The search graph for decoding is constructed as a preprocessing step using these rules. N-gram-based MT has the following drawbacks: • Only the pre-calculated orderings are hypothesized during decoding • The N-gram model can not use lexical triggers • Long distance reorderings can not be performed • Unaligned target words can not be handled • Using tuples presents a more difficult search problem than that in phrase-based SMT In this dissertation, we present a novel machine translation model based on a joint probability model, which represents translation process as a linear sequence of operations. Our model like the N-gram model uses minimal translation units, but has the ability to memorize like the phrase-based model. Unlike the “N-gram” model, our operation sequence includes not only translation but also reordering operations. The strong coupling of reordering and translation into a single generative story provides a mechanism to better restrict the position to which a word or phrase can be moved, and is able to handle short and long distance reorderings effectively. This thesis remedies the problems in phrasal MT and N-gram-based MT by making the following contributions: • We proposed a model that handles both local and long distance dependencies uniformly and effectively • Our model is able to handle discontinuous source-side units • During the decoding, we are able to remove the hard reordering constraint which is necessary in the phrase-based systems • Like the phrase-based and unlike the N-gram model, our model exhibits the ability to memorize phrases • In comparison to the N-gram-based model, our model performs search on all possible reorderings and has the ability to learn lexical triggers and apply them to unseen contexts A secondary goal of this thesis is to challenge the belief that conditional probability models work better than the joint probability models in SMT and that the source-side context is less helpful in the translation process.

Read the paper · More papers on PaperTik