A Micro-Word Based Approach for Arabic Sentiment Analysis

Fawaz S. Al–Anzi, Dia AbuZeina · 2017

Sentiment analysis of social networks data has recently received a great deal of attention. Social networks are characterized by uncommon language that is different when compared with the standard format of the language. Hence, there is a demand for effective methods to analyze the huge volume of the new word variants that quickly and daily show up in the digital and online world. In text classification, vector space model (VSM) is based on the vocabulary list (i.e. the entire training set words) while ignoring the odd words, which leads to partial loss of textual information. To address this challenge, we propose to use each two-neighboring letters of the word as a basic feature unit instead of using the word itself. That is, instead of using words in VSM, we propose a new method that is based on decomposing each word into a sequence of micro-words, each of which has only two consecutive letters. Two data collections were employed to investigate the performance. The data collections include common (i.e. standard form) and uncommon Arabic text (obtained from Instagram). For the common text, we used a corpus that contains 1,500 documents for training and 500 documents for testing. The proposed method was evaluated using latent semantic indexing (LSI) for textual features and cosine similarity measure for classification. The experimental results show promising results as the proposed method correctly classifies the testing set documents with an accuracy up to 83.6%.

Read the paper · More papers on PaperTik