Learning translation models from the Web
Jian‐Yun Nie, Chen Jiang · 2003
Query translation is-the key problem in cross-language information retrieval. It can be made by exploiting a large set of parallel texts. We describe a mining system that automatically discovers parallel Web pages on the Web. This system exploits the existing search engines, and the common characteristics in the organization of Web pages. Several large text corpora have been constructed using this system. Our experiments show that query translation using the obtained corpora can be as good as those by high-quality machine translation systems. This study shows the feasibility of building automatically a query translation system for all the active languages on the Web.