Enriching Arabic Tweets Representation based on Web Search Engine and the Rough Set Theory

Mohammed Bekkali, Issam Sahmoudi, Abdelmonaime Lachkar · 2015

Twitter is a popular micro-blogging service where users search for timely and social information. Users post short text messages called Tweets, which are limited in length. These Tweets are different from traditional documents in its shortness and sparseness. As a result, short text tends to be ambiguous without enough contextual information. To address these issues, we propose an efficient method to enrich the tweet's representation for the Arabic language using web search engine as a large and open corpus and the Rough Set Theory which is a mathematical tool to deal with vagueness and uncertainty. To assess the performance of the proposed system, a series of experiments has been conducted. The effectiveness of our system has been evaluated and compared in terms of the F1-measure using the Naïve Bayesian (NB) and the Support Vector Machine (SVM) classifiers in our Arabic Tweets Categorization System. The obtained results show that enriching the tweet's representation increases significantly the F1-measure of the Arabic tweets categorization system.

Read the paper · More papers on PaperTik