Parallel web text mining for cross-language IR

Chen Jiang, Jian‐Yun Nie · 2000

One of the approaches to cross-language information retrieval (CLIR) is based on the use of parallel texts. In this paper, we will describe a parallel text mining system called PTMiner (Parallel Text Miner) for the Web environment. We will explain the underlying mining algorithm of this system as well as its implementation using a distributed model and database technology. The resulted corpora are used as the training material for statistical translation models. Preliminary experimental results using the models for CLIR are reported. 1 Introduction Data mining, text mining and other knowledge discovering techniques have become an attractive research area in the past years. The enormous amount of information often oers potential solutions to some problems. This is the case of parallel texts that provide translation examples from a language to another. A pair of parallel texts is two such texts that are translation one for the other. In our work, the need for parallel corpora is orig...

Read the paper · More papers on PaperTik