LiSa–morphological analysis for information retrieval

Hans Hjelm, Christoph Schwarz · DSpace repository (University of Tartu) · 2006

This paper presents LiSa, a system for morphological analysis, designed to meet the needs of the Information Retrieval (IR) community.LiSa is an acronym for Linguistic and Statistical Analysis.The system is lexicon-and rule based and developed in Java.It performs lemmatization, part of speech categorization, decompounding and compound disambiguation for German, Spanish, French and English, with the other major European languages under development.The lessons learned when developing the rules for disambiguation of German compounds are also applicable to other compounding languages, such as the Nordic languages.Since compounding is much more common and far more complex in German than in the other languages currently handled by LiSa, this paper will deal mainly with German.LiSa is developed by Intrafind Software AG, on whose homepage an online demo of LiSa can be found 2 .It is used in Intrafind's iFinder and also exists as an add-on to the open source free text indexing tool Lucene 3 .

Read the paper · More papers on PaperTik