Training on the poles for review sentiment polarity classification

Michael Kranzlein, Dan Chia-Tien Lo · 2017

Support Vector Machines (SVMs) are well-known tools for the task of big data text classification. This research studies the effects of omitting non-polar training samples on the performance of a SVM-based binary text classifier. The classifier operates on a large corpus of Amazon product review text bodies, sampled from various product categories and predicts the polarity of reviews. Our results show that training on a smaller, more concise training set of only 1 and 5-star reviews offers similar performance to training on a much larger dataset that includes 1, 2, 4, and 5-star reviews, without sacrificing much in terms of precision or recall. These results hold true using both count-based and TF-IDF vectorization methodologies. The technique of training on the poles is demonstrated to offer efficient means to build a high-performing SVM-based review sentiment polarity classifier, especially in cases where labeled data may not be readily available and training time is constrained.

Read the paper · More papers on PaperTik