AraSentiment: Arabic Sentiment Analysis on Data using Machine Learning, and Transformers

Diaa Salama AbdElminaam, Abdelrahman Shorim, Mina Antoun, Habiba C. Mohamed · 2023

A natural language processing method called sentiment analysis automatically detects and measures subjective information in text or audio data. It may be used in many different situations. However, due to the complexity of the Arabic language and the scarcity of labeled audio datasets, sentiment analysis on Arabic audio data poses particular difficulties. This study suggests a fresh method for analyzing Arabic sentiment in audio data utilizing machine learning and transformers. With the Count Vectorizer and n-gram techniques serving as the feature extractors, the objective is to assess the effectiveness of several machine learning algorithms, such as AraBERT, Linear Regression, Decision Tree, Random Forest, and Naive Bayes. The suggested method seeks to achieve high accuracy in sentiment categorization by overcoming the difficulties of Arabic audio sentiment analysis. The study evaluates the proposed approach on a newly created Arabic audio dataset and compares the results with existing approaches. The obtained results in terms of Accuracy (ACC), Precision (PREC), and Recall (REC) show excellent results, with the proposed approach outperforming existing approaches. The discussion section analyzes the impact of different feature extraction methods and machine learning algorithms on the performance of the sentiment analysis model. The results show that the AraBERT model achieves the highest accuracy.

Read the paper · More papers on PaperTik