Applying Modified TF-IDF with Collocation in Classifying Disaster-Related Tweets
Gleen A. Dalaorao · International Journal of Advanced Trends in Computer Science and Engineering · 2020
Disaster-related tweets classification refers to posted tweets on twitter during the time-critical events (e.g., natural, human-made) that are group together according to pre-defined categories (e.g., donations, awareness, help, etc.).The purpose of classification is to deliver the conveyed messages to suitable authorities on time for requiring needs for immediate action.The classification works well if terms that are extracted from tweets are carefully selected or referred to as good attributes that can best label the uncategorized tweets.The Term Frequency -Inverse Document Frequency (TF-IDF) helps classifier to extract useful features.Hence, TF-IDF naturally works on distinct terms.However, a single term occasionally can be ambiguous, which means that when a separate term used for indexing, could carry numerous connotations.The distinct term can sometimes too broad, which means it does not have a discerning power to discriminate terms, for example, from the two individual terms "college" and "junior".These two terms are not adequate to differentiate "college junior" from "junior college" [1].Hence, applying traditional TF-IDF in text classification can reduce classification efficacy.Thus, a combination of terms known as collocation is introduced as an improvement to TF-IDF to boost the text classification effectiveness.This paper aims to provide an analysis of the efficacy of tweets classification by applying improved TF-IDF with collocation.This experiment utilizes tweets dataset from the CrisisLex website.The performance evaluation metrics considered are confusion matrix, precision, recall, and F1 score.The result shows that there is a favorable increase in the proposed study as compared to traditional TF-IDF through said evaluation metrics vary from 4% to 24%.The study also establishes that RandomForest consistently outperforms that two other compared classifiers.