Finding parallel texts on the web using cross-language information retrieval.

Achim Ruopp, Fei Xia · International Joint Conference on Natural Language Processing · 2008

Discovering parallel corpora on the web is a challenging task. In this paper, we use cross-language information retrieval techniques in combination with structural features to retrieve candidate page pairs from a commercial search engine. The candidate page pairs are then filtered using techniques described by Resnik and Smith (2003) to determine if they are translations. The results allow the comparison of efficiency of different parameter settings and provide an estimate for the percentage of pages that are parallel for a certain language pair.

Read the paper · More papers on PaperTik