Sentiment wEight of N-grams in Dataset (SEND): A Feature-set for Cross-domain Sentiment Classification

Atanu Dey, Mamata Jenamani, Jitesh J. Thakkar · 2017

Generation of labeled reviews is a costly and time consuming effort. Cross-domain classification reduces such effort by considering source and target dataset from two different domains. In this paper, we propose SEND - a lexicon-based feature space for sentiment n-grams that performs well compared to the existing unigram and n-gram based features typically used in same domain analysis. The features are constructed by extracting intensifiers, negations and sentiment unigrams from the dataset. We find an importance-value for each feature by computing the product of 1) the number of times the feature appears in the review and 2) the logarithmic operation of its inverse frequency in the corpus. Next, we calculate the sentiment score of the n-grams by using the individual sentiment scores of the unigrams and pre-calculated values of intensifiers, negations attached with it. These scores are multiplied with the corresponding-feature-importance- value to generate the final score of SEND features for each review. We experiment with Maximum Entropy Classifier on two benchmark datasets for cross-domain classification problem and obtain a substantial improvement in performance measure compared with state-of-the-art methods.

Read the paper · More papers on PaperTik