An affix removal stemmer for Gujarati text

Nikita P. Desai, Bijal Dalwadi · International Conference on Computing for Sustainable Global Development · 2016

In this paper, we suggest a stemmer for Gujarati language, which conflates terms by affix removal. Affix is combination of suffix and prefix. Suffix is attached after root word where as prefix is attached before it We are using suffix rules, prefix rules and substitution rules. As per our knowledge, negligible work has been done based on prefix removal in Gujarati language. We are using dictionary lookup and customized rules for Gujarati language. The dictionary is organized as red-black tree data structure. We are analyzing result on correctness of stemmer based on over-stemming and under-stemming errors. The stemmer is entirely rule-based, for which, we attained an accuracy of 96.63% with the help of specially formed suffix-stripping, substitution and prefix rules. The system developed might be useful in applications such as Information Retrieval, dictionary search, document summarization, as preprocessing module in many NLP problems, for Machine Translation, et al.

Read the paper · More papers on PaperTik