Urdu Sentiment Analysis Using Deep Learning

Imtisall Nasir, Sibghat Ullah Bazai, Muhammad Imran Ghafoor, Shah Marjan · 2024

Sentiment Analysis is a deep text analysis to determine the expressions from the written text, The Urdu language holds significant importance due to the usage of these social media platforms and content generated at these platform in Urdu. Effective sentiment analysis in Urdu can provide valuable information about social trends, customer feedback, and public opinion on different sets of topics in the Urdu context. This study uses an deep learning (ANN) Artificial Neural Network to predict and classify the sentiments in the Urdu language, labeling them into positive, negative, and neutral sentiments. To achieve the research goals, advanced preprocessing techniques were utilized such as removing stop words, emojis, emails, and characters, removing symbols, performing lemmatization to expand the meaning of text, tokenized each word by using Lexicon-based approach and ngram. The dataset are tweets dependent and have 1,140,021 tweets. These steps ensure that data is standardized and cleaned to be split into training and testing. The datasets were converted into sequence and categorized to be input into the ANN model. The ANN model was trained and achieved an accuracy of 90%, precision of 98%, and f1-score of 97%. These high values indicate that the model accurately predicted the sentiments in the Urdu text. However, during experimentation limitations were discovered due to the complexities of Urdu text, difficult script patterns, and linguistic problems that were not solved by the preprocessing technique. During analysis, the limitation was discovered due to the limited number of preprocessing techniques and tools for Urdu text. However, the study highlights the potential of neural networks, especially ANN in Urdu sentiment analysis. This study contributed significantly to Urdu natural language processing by providing robust techniques and identifying limitations to work on it as a research problem.

Read the paper · More papers on PaperTik