Learning Bilingual Word Representations by Marginalizing Alignments
Tomáš Kočiský, Karl Moritz Hermann, Phil Blunsom · 2014
We present a probabilistic model that simultaneously learns alignments and distributed representations for bilingual data.By marginalizing over word alignments the model captures a larger semantic context than prior work relying on hard alignments.The advantage of this approach is demonstrated in a cross-lingual classification task, where we outperform the prior published state of the art.