Evaluation of Tweets for Content Analysis Using Machine Learning Models

Amit Purushottam Pimpalkar, R. Jeberson Retna Raj · 2020

Now-a-days we are virtually residing in the digital networking era, and constantly sharing thoughts, opinions about things that are associated with us. As a result, social networking sites, especially Twitter, get flooded with tremendous information in the form of opinions. They are unstructured so disseminated, and a solid base has to be established such that they can be seen as useful knowledge about a specific issue. This work focuses on the evaluation of different models with respect to feature selected and dataset of user's feelings on twitter data using different pre-processing measures. The dataset is extracted from Kaggle and Twitter, pre-processing performed using NLTK, Scikit-learn and features selection, extraction done for a Bag of Words (BOW), Term Frequency (TF) and Inverse Document Frequency (IDF). For polarity identification and evaluation, five different Machine Learning (ML) algorithms were compared. The performance comparison of these algorithms in an attempt to decide which algorithm is ideally suited for the selected dataset in view of the recall, accuracy, Fl-score, and precision observed. As the evaluation results, an SVM classifier outperformed the other implemented models.

Read the paper · More papers on PaperTik