Cross-lingual document similarity

Andrej Muhič, Jan Rupnik, Primož Škraba · Information Technology Interfaces · 2012

In this paper we investigated how to compute similarities between documents written in different languages based on a weekly aligned multi-lingual collection of documents. Computing the cross-lingual similarities is based on an aligned set of basis vectors obtained by either latent semantic indexing or the k-means algorithm on an aligned multi-lingual corpus. We evaluated the methods on two data sets: Wikipedia and European Parliament Proceedings Parallel Corpus.

Read the paper · More papers on PaperTik