Tweet normalization: A knowledge based approach
Itisha Gupta, Nisheeth Joshi · 2017
Twitter Sentiment Analysis has attracted a lot of attention recently due to its promising commercial advantages. Data pre-processing is a leading and fundamental step towards sentiment analysis since it may lead to increase the correctly classified instance. The occurrence of slangs, misspelled words, hashtags, URLs, emoticons etc. in tweets make Twitter text notoriously noisy and unstructured thus sentiment analysis of Twitter text is a challenging task. So focus of this paper is to highlight pre-processing significance that would normalize and clean tweets for sentiment classification fortification. In this research work, we explore various pre-processing methods and presented a framework of the pre-processing system with a detailed description of the process which comprises of 2 segments: denoising such as removal of StopWords, URLs, username, punctuation etc. and normalization such as conversion of Non-standard words to their canonical forms. Tweets are pre-processed by normalizing elongated words (loooove to looove), misspelled words (kkk to okay), informal acronyms (rofl to rolling on the floor laughing), negation handling, emoticon replacement, etc. An evaluation is also carried out for analyzing the performance of the pre-processing system by comparison of manually pre-processed tweets and automatic pre-processed tweets and report 87.6% accuracy of proposed pre-processing method on 1000 tweets of demonetization, 90.5% accuracy on 200 tweets of iphone7 and 88.08% combined accuracy on 1200 tweets (1000 of demonetization and 200 of iphone7).