Classification of Disaster Specific Tweets - A Hybrid Approach

Siddu P. Algur, Shreya Venugopal · International Conference on Computing for Sustainable Global Development · 2021

The pervasive increase in tweeting habits of netizens has thrown open a plethora of avenues for Twitter data exploration. Users are tweeting on an innumerable number of topics with politics, entertainment, sports, and technology-related issues comprising the bulk of the tweets. Identifying tweets related to disasters - natural and manmade, among the huge repository of data posted every day on Twitter is a stimulating task. In this work, we propose a hybrid classification model to identify and distinguish the tweets attributed to disasters. An attempt is made to extract the disaster-related tweets by identifying the disaster keywords in the tweets, the tweets are then transformed into vectors using the count vectorization and Term Frequency-Inverse Document Frequency TF-IDF models. We then prepare a repository of a dataset based on unigram, bigram, and n-gram and evaluate multiple associations among the contents of a repository. Classification algorithms Naive Bayes, Logistic Regression, J48, Random Forest, and SVM are applied to evaluate their performances on our hybrid classification data model. The classification results are compared and evaluated to have an inkling about the suitability of classification with the proposed hybrid model approach.

Read the paper · More papers on PaperTik