Sarcasm classification: A novel approach by using Content Based Feature Selection Method

H. M. Keerthi Kumar, Bukahally Somashekar Harish · Procedia Computer Science · 2018

In recent decades, social media sites such as twitter, facebook, and review site produces huge number of textual information posted by many users. The user tends to express his/her sentiment in the form of sarcastic utterances. The sarcastic utterance usually shifts the polarity of text from negative to positive and likewise. The automatic classification of sarcastic utterances present in text is a very challenge task. It requires a system that can manage to detect content based text properties or features present in sarcastic utterances. In this regard, the paper propose a novel approach to classify sarcastic text using content based feature selection method. The proposed approach consists of two stage feature selection method to select most representative features. In first stage, conventional feature selection methods such as CHI-square, Information Gain (IG) and Mutual Information (MI) are used to select relevant features subset. The selected feature subset are further refined using second stage. In second stage, k-means clustering algorithm is used to select most representative feature among similar features. The selected features are classified using two classifiers Support Vector Machine (SVM) and Random Forest (RF). The proposed approach out-performance the existing methods in terms of Precision, Recall, F-measure on Amazon product review dataset.

Read the paper · More papers on PaperTik