Language Identification Strategies for Cross Language Information Retrieval.

Alessio Bosca, Luca Dini · 2010

Abstract. In our participation to the 2010 LogCLEF track we focused on the analysis of the European Library (TEL) logs and in particular we experimented with the identification of the natural language used in the queries. Language identification is in fact a key task within Cross Language Information Retrieval systems and the challenge is particularly difficult in the case of search queries where the contextual information available is scarce; function words (grammar particles highly connotative of a specific language like prepositions, pronouns, conjunctions, etc) are usually missing and the relevant presence of Named Entities can be misleading for the correct identification of the language used in the query. In order to face this challenge with acceptable performances the techniques applied should be different form the ones adopted for language guessing with more extensive and syntactically richer text fragments, like metadata or textual documents. In particular we experimented combining together different strategies: corpus based, character model based and a priori hypothesis. Since no official evaluation of the task is available we manually evaluated a sample of 100 queries and the results obtained are quite promising. Keywords: Cross-Language Information Retrieval, Language Identification, Log Analysis.

Read the paper · More papers on PaperTik