Latent Domain Word Alignment for Heterogeneous Corpora

Hoang Manh Cuong, Khalil Sima’an · 2015

This work focuses on the insensitivity of existing word alignment models to domain differences, which often yields suboptimal results on large heterogeneous data.A novel latent domain word alignment model is proposed, which induces domain-conditioned lexical and alignment statistics.We propose to train the model on a heterogeneous corpus under partial supervision, using a small number of seed samples from different domains.The seed samples allow estimating sharper, domain-conditioned word alignment statistics for sentence pairs.Our experiments show that the derived domain-conditioned statistics, once combined together, produce notable improvements both in word alignment accuracy and in translation accuracy of their resulting SMT systems.

Read the paper · More papers on PaperTik