CoLesIR at CLEF 2006: Rapid Prototyping of an N-gram-Based CLIR System
Jesús Vilares, Michael Oakes, John Irving Tait · 2006
In this our first joint participation as the CoLesIR group, our team has participated in the Portuguese monolingual ad-hoc task and in all robust ad-hoc tasks —all monolingual tasks, the English-to-German bilingual task, and the multilingual task. We have developed an n-gram model inspired by the previous work of the Johns Hopkins University Applied Physics Lab. Our approach makes generalized use of freely available resources —such as the Europarl parallel corpus, the GIZA++ wordalignment toolkit, and the Terrier retrieval platform—, and employs a new n-gram direct translation technique. This new technique takes as input previously existing aligned word lists and obtains as output aligned n-gram lists. It can also handle word translation probabilities, as in the case of statistical word alignments. This new n-gram-based approach shares the main advantages of the original proposal. This solution avoids the need for word normalization during indexing or translation, and it can also deal with out-of-vocabulary words. Since it does not rely on language-specific processing, it can be applied to very different languages, even when