In-depth Urdu Sentiment Analysis Through Multilingual BERT and Supervised Learning Approaches

Muhammad Imran Saeed, Naeem Ahmed, Danish Ali, Muhammad Ramzan, Muzamil Mohib, Kajol Bagga, Atif Ur Rahman, Ikram Majeed Khan · ICCK Transactions on Intelligent Systematics · 2024

Sentiment analysis is a crucial component of intelligent information processing systems, enabling machines to understand and categorize human opinions expressed in text. While extensively studied for high-resource languages such as English and Chinese, it remains underexplored for low-resource languages like Urdu. This paper presents an intelligent multilingual sentiment analysis framework for Urdu text by integrating supervised machine learning techniques with a transformer-based model. We manually annotated and preprocessed a dataset collected from various Urdu blog websites, categorizing sentiments into positive, neutral, and negative classes. Four machine learning classifiers—Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Naive Bayes, and Multinomial Logistic Regression (MLR)—along with the transformer-based multilingual BERT (mBERT) model were systematically evaluated. The mBERT model was fine-tuned to capture deep contextual embeddings tailored for Urdu, leveraging transfer learning from a model pre-trained on 104 languages. Experimental results demonstrate that the proposed intelligent framework significantly outperforms traditional classifiers, achieving an accuracy of 96.5% on the test set. This study highlights the effectiveness of transfer learning and deep contextual models in building robust intelligent systems for low-resource language processing, contributing to the advancement of inclusive and systematic intelligence in natural language understanding.

Read the paper · More papers on PaperTik