BanglaLem: A Transformer-based Bangla Lemmatizer with an Enhanced Dataset

Md Fuadul Islam, Jakir Hasan, Md. Ashikul Islam, Prato Dewan, M. Shahidur Rahman · Systems and Soft Computing · 2025

Lemmatization plays a crucial role in various natural language processing (NLP) tasks, such as information retrieval, sentiment analysis, text summarization, and text classification. However, Bangla lemmatization remains particularly challenging due to the language’s rich morphology and high inflectional complexity. Existing open-access datasets for Bangla lemmatization are limited in size, with the largest containing only 22353 unique inflected words, which constrains the effectiveness of data-driven neural models. To address this limitation, we introduce a novel dataset, BanglaLem, comprising 96040 frequently used inflected words. This dataset has been carefully curated and annotated through a rigorous selection process to enhance the accuracy and efficiency of Bangla lemmatization. Furthermore, we propose a transformer-based approach to lemmatization and evaluate the performance of various pre-trained and trained from-scratch transformer models on this dataset. Among these, the BanglaT5 model achieved the highest exact match accuracy of 94.42% on the test set. The BanglaLem dataset is publicly accessible via the following link . • Developed a transformer-based Bangla lemmatizer with high accuracy. • Created a novel dataset of 96040 inflected Bangla words. • Addressed scarcity of large, open-access Bangla lemmatization datasets. • Achieved 94.42% exact match accuracy with the BanglaT5 model on testing data.

Read the paper · More papers on PaperTik