HAL: Challenging Three Key Aspects of IBM-style Statistical Machine Translation

Christer Samuelsson · Conference of the Association for Machine Translation in the Americas · 2012

The IBM schemes use weighted cooccurrence counts to iteratively improve translation and alignment probability estimates. We argue that: 1) these cooccurrence counts should be combined differently to capture word correlation; 2) alignment probabilities adopt predictable distributions; and 3) consequently, no iteration is needed. This applies equally well to word-based and phrase-based approaches. The resulting scheme, dubbed HAL, outperforms the IBM scheme in experiments.

Read the paper · More papers on PaperTik