A topic-based cross-language retrieval model with PLSA and TF-IDF

Zhuo-Liang Huo, Junfeng Wu, Yao Lu, Changzhou Li · 2018

Cross-language document retrieval is an important research topic in information retrieval. This paper proposes a cross-language retrieval model, which can retrieve English documents for a Chinese query. The model includes two parts: topic-processing and online retrieval. The topic-processing could be executed once the library of literature is updated. It extracts topics from semantics of texts and then clusters documents of different languages with the extracted topics. Three crucial algorithms are adopted in topic-processing: Probability Latent Semantic Analysis (PLSA) analyzes the topics of documents, Term Frequency-Inverse Document Frequency (TFIDF) identifies keywords for the topics, and word co-occurrence method with dictionary-based translation generates the association among words of different languages. Practical retrieval runs whenever a retrieval request is submitted by a user. It achieves our goal of cross-languages retrieval by analyzing the correlation between query terms and topics, and using document correlation results from topic-processing to find appropriate document results.

Read the paper · More papers on PaperTik