Combination of Statistical Word Alignments Based on Multiple Preprocessing Schemes
Jakob Elming, Nizar Y. Habash, Josep Crego · The MIT Press eBooks · 2008
This chapter presents an approach to using multiple preprocessing (tokenization) schemes to improve statistical word alignments. In this approach, the text to align is tokenized before statistical alignment, and then remapped to its original form afterwards. Multiple tokenizations yield multiple remappings (remapped alignments), which are then combined using supervised machine learning. The remapping strategy improves alignment correctness by itself. The combination of multiple remappings also improves measurably over a commonly used state-of-the-art baseline. A relative reduction of alignment error rate of about 38% is obtained on a blind test set.