NTCIR-2 Chinese, Cross Language Retrieval Experiments Using PIRCS.
Kui Lam Kwok · 2001
We participated in the monolingual Chinese and English-Chinese cross language retrieval track using our PIRCS retrieval system. Employing the query translation approach for crosslingual IR, two methods of translation were tried: MT software, and dictionary lookup followed with disambiguation techniques. Retrieval lists from the two methods were combined to form the final result. Pseudo-relevance feedback was used, but no pre-translation expansion or collection enrichment was employed. Runs for all query lengths were submitted. Short-word with character representation was used for both documents and queries. Using the ‘rigid ’ criteria for evaluation, both VS (very short) and SO (short) queries gave monolingual mean average precision at over 0.6. LO (long) query type surprisingly was worse by about 5 % at over 0.57. TI queries (title only of a few words) returned a good 0.46 mean average precision. Cross language retrievals perform at between 77 to 83 % of monolingual for the longer query types, and only at 55 % for the TI queries. A prominent factor in cross language retrieval failure is unknown word (such as proper noun) translation. This is particularly acute with TI queries of a few words. We show that it may be overcome to some extent by using longer queries that can provide more redundancy and better context for translations to hedge for errors. Crosslingual IR can also be efficiently performed, often with improved results, by mixting both translations as one single query.