Rough Set Techniques for Text Classification and Sentiment Analysis in Social Media

G. K. Panda, Jayanta Mondal, Sriram Vihar · 2015

Abstract: Sentiment Analysis (SA) is an ongoing research in the field of text mining and classification. SA finds a computational domain from opinions and subjectivity of text data in online social media. Sentiments are inherited in the form of simple lexicons with symbols and texts having noise of irregular texts in complex forms. It is also seen that the high dimensional growth of lexical blends used by online users while expressing or responding their responses. These blends differ according to demographics and on the context of topics. The simplest approach to get rid of the noise data, adapted by number of studies is by simply removing the irregular lexicons, stopwords, emoticons and lexical blends. This paper investigates such effectiveness in sentiment classification. We assess the impact of rough set approach in classification of universe and apply to the raw datasets. Our earlier study on covering based approximation of classifications outperforms the general classification of universe. We apply roughest based classification process using MATLAB functions to the raw dataset before data pre-processing. Our results show that precompiled roughest classification has better accuracy and outperforms than some of earlier studies.

Read the paper · More papers on PaperTik