Sentiment Analysis of Libyan Dialect Using Machine Learning with Stemming and Stop-words Removal

Abdullah Habberrih, Mustafa Ali Abuzaraida · 2024

This study evaluates the impact of using Stemming and Stop-words removal techniques on machine learning classifiers, Support Vector Machine (SVM) and Logistic Regression (LR), in detecting sentiment from Libyan dialect poetry. The lack of Arabic Natural Language Process resources has made sentiment analysis for Arabic a challenging task compared to other languages. A secondary dataset was used and two experiments were conducted, with the first exploring the use of Stemming with Stop-words removal techniques and the second investigating the impact of using Stemming alone. Other preprocessing techniques were applied alongside TF-IDF with a combination of Unigrams and Trigrams during feature extraction. The results show that the Stop-words removal technique may have a negative impact on classifier performance. SVM outperformed LR in both experiments, achieving an accuracy of 71.63%, while LR achieved 70.92% in the second experiment. This study's accuracy outperformed previous research on the topic, achieving 71.63%, compared to 69% in earlier studies.

Read the paper · More papers on PaperTik