Distortion constraints in statistical machine translation

DAN I. MOLDOVAN, Marian Olteanu · 2007

Statistical machine translation (SMT) offers many benefits over rule-based and example-based machine translation especially ease to train and robustness. It represents state-of-the-art in machine translation (MT) but it has to deal with certain issues that are not trivial to solve in a statistical framework: correct distortion, correct agreement, morphology issues, etc. In order to solve distortion issues, different models were proposed: syntax-based MT, enhanced distortion models, clause restructuring. All of these define more complex distortion models than monotonic distortion models (which favor lack of word reordering). But these proposed models don't completely solve the issue of easily adding linguistic knowledge into the MT decoder. The dissertation proposed a model designed to augment SMT models with linguistic knowledge, either in a rule-based fashion or in a probabilistic fashion. The theoretical framework proposed in this dissertation was implemented in Phramer statistical phrase-based decoder and tested using various levels of knowledge—surface (only punctuation), Part-of-Speech, constituency parse trees. The experimental results showed improvement in the quality of the translation for language pairs that were tested (English-to-French, English-to-Spanish, French-to-English, Arabic-to-English). The improvement was observed in automatic metrics (BLEU score, Word Error Rate and Position independent word Error Rate) and manually, by inspecting the differences in translation. As expected, the differences were less in the lexical choice and more in word ordering that can be observed in the precision of higher-order n-grams that is used in the computation of the BLEU metric.

Read the paper · More papers on PaperTik