Named Entity Recognition for Telugu and Kannada on Naamapadam Dataset Using Various Machine Learning and Deep Learning Algorithms

Kuruva Divya Sree, Gottam Geethika, Aarathi Rajagopalan Nair, Deepa Gupta, G. S. Veena · Procedia Computer Science · 2025

Named Entity Recognition is a crucial task in Natural Language Processing. It involves detecting and categorizing named entities such as individuals, places, and organizations within a text. This process is vital for transforming un-structured text into structured information, which is useful for various applications like information retrieval, question answering, and text summarization. The proposed study investigates bilingual NER for Telugu and Kannada languages using pre-trained language models and embeddings for contextual vectorization. Utilizing the Naamapadam dataset, the research employs IndicFT, MuRIL and IndicBERT embeddings along with machine learning and deep learning algorithms. Notably, utilized only 10% of the considered dataset and achieved results that surpass those reported in the base paper. Among the models tested, MuRIL combined with XGBoost achieved the best performance, with an F1 score of 89.00 for Telugu and 87.00 for Kannada. Performance evaluation demonstrates varying effectiveness across techniques, shedding light on optimal strategies for NER in under resourced languages. These findings contribute to advancing NLP techniques for Indian languages, facilitating broader accessibility and usability in diverse linguistic contexts.

Read the paper · More papers on PaperTik