Comparison of a Semi-Automatically Generated and a Manually Created Linguistic Rule-Based Stemmer for Polish by a Polish Native Speaker
Nikitas Ν. Karanikolas, Mateusz Zywczak · 2024
Usually, Stemmers are manually created by native target language speakers, on the basis of rules for suffix stripping or replacements. In such cases linguistic knowledge of the target language is a requirement. There are also approaches for automatic creation of stemmers. The later approaches are usually based on statistical processing of text corpuses to calculate the occurrences and other metrics for stems and affixes. In such cases linguistic knowledge of the target language is not needed. We have invented a semi-automatic methodology (and we have made an equivalent implementation) for Stemmer generation that neither linguistic knowledge of target language nor statistical processing of corpuses is needed. In our approach only Information Retrieval expertise is needed. Here we are evaluating our approach against the rule-based (linguistic-knowledge-based) approach. To do so we evaluate two stemmers by a native speaker of the target (Polish) language. The first Stemmer is rule-based (it incorporates linguistic knowledge of Polish language) and the second one is based on our methodology (only the experience of Information Retrieval Experts is incorporated). The results are interesting and hard to wait for, they are in favor of our methodology.