Analysis of Multimodal Social Media Data Utilizing VIT Base 16 and GPT-2 for Disaster Response

Shilpa Gite, Shruti Patil, Biswajeet Pradhan, Madhuri Yadav, Sneha Basak, Arundarasi Rajendra, Abdullah M. Alamri, Kaustubh Raykar, Ketan V. Kotecha · Arabian Journal for Science and Engineering · 2025

Abstract Multimedia systems, such as social media platforms, play a crucial role in disseminating vital information during calamities. This information is shared in various formats such as images, text, videos, audio, etc. Therefore, it becomes important to have a system that can identify multimodal data to classify relevant information. This paper proposes a new age classification method for multimodal data using advanced and improved transformer models, such as Vision Transformer and Generative Pre-trained Transformer 2, for image and text classification, respectively. These models were combined using an ensemble model (Random Forest Classifier), achieving an accuracy of 84.66% on the multimodal data. Furthermore, the proposed model demonstrates higher prediction accuracy compared to traditional Convolutional Neural Network (CNN) models which have an accuracy of 71.43%, exceeding it by 13.23%. A comparison with convolutional models is conducted to underscore the advantages of transformer models and to substantiate the necessity of the experiment. Our proposed classification model using Vision Transformer and GPT-2, along with an ensemble model, can be replicated by researchers in disaster management, humanitarian aid organizations, and social media platforms looking to filter and prioritize information during emergencies.

Read the paper · More papers on PaperTik