Choosing The Most Optimum Text Preprocessing Method for Sentiment Analysis: Case:iPhone Tweets
Fero Resyanto, Yuliant Sibaroni, Ade Romadhony · 2019 Fourth International Conference on Informatics and Computing (ICIC) · 2019
Preprocessing is the initial stage in several text processing tasks, including sentiment analysis. Preprocessing is an important step in sentiment analysis because it could affect the result accuracy significantly. However, previous studies on preprocessing that discussed the selection of preprocessing methods were rarely conducted. In this study, we analyze the effect of preprocessing methods on sentiment analysis task. We performed the sentiment analysis as a classification on product opinion, whether the sentiment is positive or negative. We conducted an experiment using Tweets that talk about iPhone. We observed seven different preprocessing methods and the combination of it. The preprocessing methods are: casefolding, expressive lengthening, emoticons handling, removing URLs, slang handling, punctuations handling, stopwords removal and stemming. The results show that a combination of five methods: URL removal, emoticon handling, case folding, expressive lengthening and stemming is the most optimum method with an accuracy of 70.88% on sentiment analysis.