Unsupervised Word Alignment Using Frequency Constraint in Posterior Regularized EM

Hidetaka Kamigaito, Taro Watanabe, Hiroya Takamura, Manabu Okumura · 2014

Generative word alignment models, such as IBM Models, are restricted to oneto-many alignment, and cannot explicitly represent many-to-many relationships in a bilingual text.The problem is partially solved either by introducing heuristics or by agreement constraints such that two directional word alignments agree with each other.In this paper, we focus on the posterior regularization framework (Ganchev et al., 2010) that can force two directional word alignment models to agree with each other during training, and propose new constraints that can take into account the difference between function words and content words.Experimental results on French-to-English and Japanese-to-English alignment tasks show statistically significant gains over the previous posterior regularization baseline.We also observed gains in Japanese-to-English translation tasks, which prove the effectiveness of our methods under grammatically different language pairs.

Read the paper · More papers on PaperTik