Using Linguistic Tools and Resources in Cross-Language Retrieval
Carol Peters, Eugenio Picchi · 1997
A system to process bilingual/multilingual text corpora is described. Thesystem includes components for crosslanguage querying on parallel (ietranslation equivalent) and comparable (ie domain-specific) collections oftexts in more than one language. Both sets of procedures are dependent on lexical resources (bilingual lexicaldatabases) and linguistic tools (morphological procedures). The system was originally designed to meet the requirements ofvarious types of contrastive language studies. However, we are now studyingapplications to cross-language retrieval. Background In the last few years, natural language processing (NLP) techniques andtools have been incorporated into information retrieval (IR) systems withvarying degrees of success (Smeaton 1992). The recent emergence of thefield of CrossLanguage Information Retrieval as an independent area of interest has clearly reinforced this trend. In order to be successful, cross-language applicationsfrequently need access to methodologies ...