An Arabic Twitter Corpus for Subjectivity and Sentiment Analysis

Eshrag Ali Refaee, Verena Rieser · 2014

We present a newly collected data set of 8,868 gold-standard annotated Arabic twitter feeds.The corpus is manually labelled for subjectivity and sentiment analysis (SSA) (κ = 0.816).In addition, the corpus is annotated with a variety of linguistically motivated feature-sets that have previously shown positive impact on classification performance.The paper highlights issues posed by twitter as a genre, such as a mixture of language varieties and topic-shifts.Our next step is to extend the current corpus, using online semi-supervised learning.A first sub-corpus will be released via the ELRA repository as part of this submission.

Read the paper · More papers on PaperTik