NTCIR-2 ECIR Experiments at Maryland: Comparing Pirkola's Structured Queries and Balanced Translation.
Douglas W. Oard, Jian‐qiang Wang · NTCIR · 2001
Pirkola’s structured queries have been shown to perform well for word-based cross-language information retrieval in European languages, but in monolingual Chinese retrieval experiments it is often found that character bigrams perform as well as, and sometimes better than, automatically segmented words. During the Mandarin-English Information (MEI) project at the Johns Hopkins Summer 2000 Workshop, Pirkola’s structured queries were compared with an alternative technique known as balanced translation. The results suggested that balanced translation coupled with post-translation character bigram resegmentation could outperform Pirkola’s word-based technique. The NTCIR-2 English/Chinese Information Retrieval (ECIR) evaluation provided the opportunity to replicate this experiment on a far larger collection. The results show that on the ECIR collection, Pirkola’s structured queries outperform balanced translation, even when post-translation character bigram resegmentation was used. This paper contrasts the MEI results with Maryland’s ECIR experiments and identifies some possible causes for the observed differences.