Learning Cross-lingual Representations with Matrix Factorization

Hanan Aldarmaki, Mona Diab · 2016

We present a matrix factorization model for learning cross-lingual representations.Using sentence-aligned corpora, the proposed model learns distributed representations by factoring the given data into language-dependent factors and one shared factor.Moreover, the model can quickly learn shared representations for more than two languages without undermining the quality of the monolingual components.The model achieves an accuracy of 88% on English to German cross-lingual document classification, and 0.8 Pearson correlation on Spanish-English cross-lingual semantic textual similarity.While the results do not beat state-of-the-art performance in these tasks, we show that the crosslingual models are at least as good as their monolingual counterparts.

Read the paper · More papers on PaperTik