Bilingual Tests with Swedish, Finnish and German Queries.

Turid Hedlund, Heikki Keskustalo, Ari Pirkola, Mikko Sepponen, Kalervo Järvelin · CLEF (Working Notes) · 2000

We used a dictionary-based approach, and performed tests in the bilingual track with three language pairs, i.e., Swedish – English (Swe-Eng), Finnish – English (Fin-Eng), and German – English (Ger-Eng). All the source languages are compound languages, i.e., languages rich in compound words. A compound word refers to a multi-word expression where the component words are written together. Our main efforts were to develop techniques for the processing of compounds, to study different types of compound languages, and to study the effects query structuring in different languages. We designed and implemented a method for automated query construction in FIN SWE GER -> ENG. The goal of this process is to extract automatically topical information from sentences written in one of the source languages (FIN, SWE, GER) and to create a target language (ENG) query. The resulting query may be either structured or unstructured. Introduction NLP-techniques have been tested for IR and CLIR for several years. The point of view has been that linguistically motivated indexing would enable the catching of sense in text and in queries differently from the non-linguistic methods used in IR, for example weighting based on word occurrences. Traditional NLP-techniques are extended also to the sub-word level, i.e., morphological decomposition and stemming (Sparck Jones 1999). So far, any great success in increasing the quality of retrieval result due to these techniques have not been reported, compared to statistical methods. The language dependent linguistic features important to IR and CLIR are, for example, the number of homographic word forms, the way to treat compounds and gender features. The main problems associated with dictionary-based CLIR are 1) phrase identification and translation, 2) source language ambiguity, 3) translation ambiguity, 4) the coverage of dictionaries, 5) the processing of inflected words, and 6) untranslatable keys, in particular proper names spelled differently in different languages (Pirkola et al. 2000) Our approach to solve the general problems for bilingual CLIR is based on 1) normalisation in indexing, 2) stopword lists, 3) normalisation of topic words, 4) splitting of compounds, 5) recognition of the right components, 6) phrase composition in target language, 7) bilingual dictionaries, and 8) structured queries. All the source languages we use, Swedish, Finnish and German are languages rich in compounds, therefore it is essential to develop techniques for the processing of compounds. Our other main interest is to compare structured and unstructured queries to solve the ambiguity problem with CLIR. We used a model developed and tested for Finnish English CLIR by Pirkola (1998; Pirkola et al. 1999).

Read the paper · More papers on PaperTik