A Feature-rich Supervised Word Alignment Model for Phrase-based Statistical Machine Translation.
Chooi-Ling Goh, Eiichiro Sumita · 2009
Word alignment plays an important role in statistical machine translation (SMT) systems. The output of word alignment can be used to build a phrase table, which is the core model in the decoding of new sentences. Most current SMT systems use GIZA++, a generative model, to automatically align words from sentence-aligned parallel corpora. GIZA++ works well when large sentence-aligned corpora are used. However, it is difficult to encode syntactic and lexical features useful for handling sparse data and unseen words, such as POS tags, affixes, lemmas, etc., using generative models. A discriminative model such as conditional random fields (CRF) can solve this problem. We treat word alignment as a labelling problem, and encode the syntactic, lexical, and contextual features. Our experiments were conducted using a 35K Chinese-English hand-aligned corpus. Our model gives better word alignment results than GIZA++ by 7 % AER. Finally, we also prove that 2 % higher BLEU score can be obtained with phrase-based SMT systems when our alignment models are used.