Boosting Cross-Language Retrieval by Learning Bilingual Phrase Associations from Relevance Rankings
Artem Sokokov, Laura Jehl, Felix Hieber, Stefan Riezler · 2013
We present an approach to learning bilingual n-gram correspondences from relevance rankings of English documents for Japanese queries.We show that directly optimizing cross-lingual rankings rivals and complements machine translation-based cross-language information retrieval (CLIR).We propose an efficient boosting algorithm that deals with very large cross-product spaces of word correspondences.We show in an experimental evaluation on patent prior art search that our approach, and in particular a consensus-based combination of boosting and translation-based approaches, yields substantial improvements in CLIR performance.Our training and test data are made publicly available.