Transfer Learning for Marathi Named Entity Recognition

Akash Misal, Yashodhara V. Haribhakta · 2022

Named entity recognition (NER) is an essential Natural Language Processing (NLP) task to find key entities from sentences. In this paper, NewsCorpus Marathi monolingual data set is created, which contains 900 sentences and 7000 tokens. Dataset is scraped from different Marathi news sources available on the internet. Paper analyzed traditional NER and POS (Part of Speech) tagging approaches for the Marathi Language. Marathi is one of the regional languages in India. Paper proposed a Pre-trained Model Approach of transfer learning for Marathi NER. IndicBERT, mBERT, and XLM-Roberta are BERT-based pretrained models that produce state-of-the-art results on a downstream task such as Named Entity Recognition. BERT model is specifically trained on Wikipedia and Google’s BooksCorpus. Paper comparing the transfer learning approach with traditional machine learning algorithms. Transfer learning-based models perform well than the conventional approach. IndicBERT is performing better among other BERT-based models on our Corpus.

Read the paper · More papers on PaperTik