Stemming techniques for resource-poor languages: a review of methods, challenges, and applications
Sanjiban Sekhar Roy, Mohd Anas, Saravanakumar Kandasamy · Frontiers in Artificial Intelligence · 2026
The process of stemming is a fundamental part of pre-processing in NLP systems in resource-poor and morphologically complex languages in which linguistic data and annotated corpora are limited. This survey presents a a systematic review of stemming techniques across Afro-Asiatic, Indo-Aryan, Turkic, and Uralic language families, with explicit acknowledgment of coverage limitation and it covers rule-based, statistical, unsupervised, hybrid, and emerging neural-assisted approaches. The evolution of stemming is analyzed in the broader context of NLP paradigms, from early suffix-stripping algorithms to modern sub-word-aware and transformer-influenced models. Special emphasis has been given on resource-poor languages such as Arabic, Urdu, Bengali, Assamese, and other low-resource Indo-Aryan, Afro-Asiatic, and agglutinative languages. For these languages performing explicit stemming remains a critical task despite advances in deep learning. This study further reviews application-driven impacts of stemming in information retrieval, text classification, sentiment analysis, topic modeling, and semantic similarity estimation. A unified comparison framework is presented which has incorporated benchmark datasets, shared-task resources, and multi-dimensional evaluation metrics, including precision, recall, F-measure, MAP, NMI, and stemming-specific error indices. Using careful comparative evaluation, this study points out the strengths and weaknesses of current approaches and provides insights into potential areas for future investigation and developments in language-sensitive stemming systems.