Multi-language text indexing for internet retrieval

Martin Wechsler, Páraic Sheridan, Peter Scháuble · 1997

We address here the issues associated with indexing multilingual collections of information, as is found for example on the internet. We examine in particular the task of language identification and the use of stemming algorithms for several European languages. We also present the lessons we have learned from our experience in using the SPIDER information retrieval system as a search engine over the intranet of the ETH Zurich; a multilingual intranet which contains documents in English, French, German and Italian.

Read the paper · More papers on PaperTik