Marathi Social Media Opinion Mining using XLM-R
Naitik Rathod, Nishit Mistry, Dhruv Talati, Manan Parikh, Aniket Kore, Pratik Kanani · 2022 International Conference on Applied Artificial Intelligence and Computing (ICAAIC) · 2022
Sentiment Analysis is one of the most important tasks for any language and a very important domain in Natural Language Processing which has shown remarkable progress recently. Popular and widely used languages like English, Russian and Spanish have a great availability of language models for these tasks and widely available datasets too. But the research in Low Resource Languages like Hindi and Marathi is far behind. The Marathi language is one of the languages mentioned in the 8thschedule of the constitution of India and is the third most spoken language in India which is mainly used in the Deccan region which includes Maharashtra and Goa. There has been low research for sentiment analysis approaches based on the Marathi text. Therefore, this project proposes use of XLM-RoBERTa (XLM-R) models that can be used for the opinion mining of the social media Marathi texts without using any translations. Not using translations will not only get better results but also an error free model trained over the target language only. The multilingual model XLM-R and its versions will be put under training over the Marathi tweets dataset after tokenizing using RoBERTa tokenizer for the purpose of opinion mining and classification. Authors aim at presenting the results of multiple XLM-R models over the Marathi tweets dataset for the task of opinion mining.