Comprehensive Approach to Dataset Creation for Sentiment Analysis in Malayalam

R. Anitha, R R Rajeev, Meharuniza Nazeem, S Navaneeth, K. S. Anil Kumar · 2024

Sentiment analysis task provides the unique challenge of analyzing a fraction of human emotion from text, Malayalam which is already a computationally complex language due to its high morphological features and agglutination further enhances these challenges. A dataset of 22,449 samples was prepared from surveys and social media that have been classified as positive, negative, or neutral. To study the applicability of the dataset a Machine learning-based sentiment analysis was carried out. TF-IDF was used for word vectorization and a Support Vector Machine (SVM) classifier with a Radial Basis Function (RBF) kernel and a Random Forest classifier was used. With an accuracy of $85.07 \%$, the RF classifier outperformed the SVM model, which only managed $59.9 \%$. The study outlines potential improvement by expanding the dataset to better represent the low-resource language of Malayalam.

Read the paper · More papers on PaperTik