TweetSentKW: A corpus of multi-label emotion analysis for Kuwaiti Arabic tweets

Eiman Alsharhan, Eisa Al Nashmi, Allan Ramsay · مجلة دراسات الخليج والجزيرة العربية. · 2024

Objectives: Arabic language is primarily represented in two varieties: Modern Standard Arabic (MSA) and Dialectal Arabic (DA). With the advent of social media, there has been a shift from the predominant use of MSA in writing to the incorporation of DA, thereby generating extensive resources for dialectal text studies. Kuwaiti Arabic (KA), a sub-variety of the Gulf dialect and one of the five principal Arabic dialects, differs significantly from MSA in all linguistic aspects. KA is an under-resourced language with a notable deficiency in language resources. The development of emotion classification tools relies heavily on the availability of resources such as annotated corpora. This study introduces TweetSentKW, a multi-label emotion annotated corpus for KA. Method: TweetSentKW was developed by collecting tweets and selecting relevant emotional classes for the annotation process. Each tweet was annotated by three independent annotators. Results: The TweetSentKW corpus comprises 40,000 manually labeled tweets across various topics. Besides constructing the corpus, this study provides a comprehensive analysis of annotator behavior and the co-occurrences of emotions. The corpus is anticipated to significantly contribute to sentiment analysis research, a crucial method for gauging public opinion. Conclusion: The widespread use of social media platforms, such as Twitter, has led to continuous and uninhibited public expression of opinions on diverse issues. The public and archived nature of these opinions presents a rich opportunity for researchers to analyze and understand public sentiment and perspectives.

Read the paper · More papers on PaperTik