English-Japanese Cross-lingual Query Expansion Using Random Indexing of Aligned Bilingual Text Data

Magnus Sahlgren, Preben Hansen, Jussi Karlgren · 2002

Vector-space techniques can be used for extracting semantically similar words from the co-occurrence statistics of words in large text data. In this paper, we report on experiments with using the Random Indexing vector-space technique for extracting a cross-lingual thesaurus from aligned English-Japanese bilingual data. The cross-lingual thesaurus has been used for automatic cross-lingual query expansion in the NTCIR patent retrieval task.

Read the paper · More papers on PaperTik