Transformer-based Models for Language Identification: A Comparative Study

Bharathi Mohan G, R Prasanna Kumar, R Elakkiya, R. Venkatakrishnan, Harrieni Shankar, Y Sree Harshitha, K. Harini, Mahati Reddy · 2023

Language Identification (LI) is a crucial task in natural language processing (NLP) with significant implications for various NLP applications, including machine translation, sentiment analysis, and text summarization. In this paper, we present a novel LI approach that harnesses the power of transformers, which are state-of-the-art models in NLP. We conduct a comprehensive comparative study between the DistilBERT model, ALBERT model and the XLM-RoBERTa model for LI. Our findings reveal that the DistilBERT model, with its lightweight design, offers comparable accuracy with reduced computational requirements. ALBERT model is designed to achieve parameter reduction by employing parameter-sharing techniques and factorized embedding parameterization. Further-more, we observe that XLM-RoBERTa, a multilingual variant of RoBERTa, surpasses both ALBERT and DistilBERT in handling diverse LI tasks.

Read the paper · More papers on PaperTik