Indian Language Analysis with XLM-RoBERTa: Enhancing Parts of Speech Tagging for Effective Natural Language Preprocessing

K Krishna Jayanth, G Bharathi Mohan, R Prasanna Kumar · 2023

This research examines the effectiveness of XLMRoBERTa, a potent deep learning architecture based on transformers, for POS labeling in Indian languages. Specifically, it concentrates on the UCD dataset(Telugu,Hindi,Tamil) and Kannada news dataset known for their complex morphological structure. By fine-tuning the pre-trained XLM-RoBERTa model, nuanced patterns between words and their POS identifiers are captured for Telugu(91%),Hindi(93%),Tamil(90%), Kannada(91%). The study provides a comprehensive evaluation, demonstrating that XLM-RoBERTa obtains a combined accuracy score of 93% for all languages. These findings emphasize the superior performance of XLM-RoBERTa in accurately allocating POS tags, notably for morphologically complex Indian languages. The research contributes valuable insights for NLP practitioners, demonstrating XLM-RoBERTa as a robust and reliable model for POS labeling and opening avenues for further advancements in deep learning architectures for NLP tasks.

Read the paper · More papers on PaperTik