Disaster Tweet Classification using LSTM: A Comparative Study of Imbalanced Data Handling Techniques
Hanif Al Irsyad, Arif Dwi Laksito · 2023
This research explores the impact of three balancing methods (SMOTE, Random Over-Sampling, and AdaSyn) on LSTM model accuracy in disaster tweet text classification. It aims to identify the most effective data balancing technique to enhance the LSTM algorithm’s performance. The study compares LSTM model accuracy using different balancing methods and determines the optimal approach for precise disaster tweet classification. The combination of these approaches enables effective processing of sequential data, enhancing the overall analysis of tweets in disaster scenarios. The LSTM algorithm, without balancing, achieved 81.6% average accuracy across ten tests on disaster tweet data. When combined with balancing methods, accuracy ranged from 81% to 81.1%. These findings demonstrate the LSTM algorithm’s strong performance in classifying disaster tweets, with a slight accuracy reduction when using balancing techniques. The research concludes that using the LSTM algorithm for text classification of disaster tweets without balancing the dataset achieves higher accuracy. However, further research is required for algorithm enhancement and performance improvement.