A combination of low-level light stemming and support vector machines for the classification of Arabic opinions
Walid Cherif, Abdellah Madani, Mohamed Kissi · 2016
Recent years have brought the burst of volume of shared opinionated texts across the internet. Every day, a tremendous number of comments and reviews towards different aspects of our lives is generated through social networks and other websites. A large portion of these data is written in Arabic which is the fifth most used language on internet and is one of the six official languages of the United Nations. However, the templatic morphology of Arabic language makes it a difficult language in the fields of information retrieval. In light of the scarcity of Arabic opinion mining systems, this paper introduces a complete approach to prepare, model and classify Arabic comments and reviews. It is conducted using a dataset consisting of 500 comments collected from TripAdvisor Website which fall into four scaled classes of opinions. For this, a new stemming technique has been defined. Stemmed opinions are then quantified according to a set of mathematical functions, and classified by different machine learning techniques. Support vector machines have yielded the highest f-measure. Finally, the impact of each step on the overall performance has been evaluated.