Tokenization and Stemming of Limbu Language
Abigail Rai, Samarjeet Borah · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
Significant issues in tokenization and stemming in natural language processing are addressed in this work, focusing primarily on the Limbu language. Two essential preprocessing procedures that work to normalize words by condensing them into compact content and their origin are stemming and tokenization. We introduce a novel stemmer for the Limbu language, which achieves 93.5% accuracy rates using the given word bank. The stemmer is preceded by tokenization with a 100% accuracy rate on the current available corpus. The Limbu language presents peculiar issues due to its constrained computational resources and composite nature. This work is a first step toward overcoming the lack of necessary resources for effective stemming in Limbu language computational work. Advanced algorithms with morphologically structured rules especially chosen for the Limbu language are used in stemming techniques to increase accuracy and speed. The discoveries significantly improve natural language processing activities, providing strong tools for search engines, sentiment analysis, and automatic translation systems, as well as marking a first step in computational development for the Limbu language.