Cross-Lingual Information Retrieval based on LSI with Multiple Word Spaces.

Tatsunori Mori, Tomoharu Kokubu, T. Tanaka · 2001

In this paper, we report the utilization of a large-scaled bilingual corpus in Cross-Language Latent Semantic Indexing(CL-LSI). When we construct one monolithic word space with a large-scaled corpus, we encounter problems such as the increase in ambigu-ity of word translation, the difficulty in singular value decomposition, which is the important process in LSI. In order to cope with the problems, we introduce the method in which the large bilingual corpus is divided into smaller sub-corpora according to the similarity among documents in it, and from each of them one word sub-space is created. By placing each docu-ment in the word sub-space, which is made from the sub-corpus most similar to the document, ambiguity of translation is expected to decrease. In the search-ing process, the query is placed into every word sub-spaces, and similarity between the query and the doc-uments are calculated. 1

Read the paper · More papers on PaperTik