Annotated Data Augmentation for Arabic Sentiment Analysis using Semi-Supervised GANs

Asma AlNashash, Ahmad T. Al‐Taani, Saleh M. Abu-Soud · Research Square · 2022

Abstract With the increasing number of web users, a tremendous amount of data is being generated in an enormous and dynamic way. In this regard, analyzing the sentiment of the user’s generated data is becoming essential to reveal valuable insights. However, analyzing this massive amount of data requires a lot of manual effort. Therefore, sentiment analysis systems automate the analysis of the generated text which helps in extracting people’s opinions and attitudes. The great success of sentiment analysis systems is due to the availability of a copiously annotated corpus. Unfortunately, Arabic Language does not possess such an abundance of annotated data. Obtaining annotated sentiment corpus is an expensive and challenging task due to the huge amount of corpus required to train efficient sentiment analyzers. The purpose of this research is to tackle the problems of Sentiment Analysis for Arabic language and the inadequate annotated data available for this task. Owing to the limited resources for Arabic sentiment analysis, this work is motivated to investigate data augmentation using GAN. GANs showed breakthrough results in data augmentation due to their capabilities of generating realistic data samples. In this work, we propose using Semi-Supervised GANs for Arabic sentiment analysis to minimize the number of annotated data required for this task. As a result, we were able to prove that GAN can be used to augment realistic textual data samples that can improve the performance of the sentiment analysis task. Results also showed that Semi-Supervised GAN can drastically reduce the requirement of manually annotated samples. In this context, our model successfully excels the task of Arabic Sentiment Analysis with only 10–30% of dataset being manually labeled.

Read the paper · More papers on PaperTik