Information Management using Natural Language Processing: A COVID-19 Case Study
Reshma Kar, Venu Devadi, Gulhan Bizel, Jyotishman Pathak, Braja Gopal Patra · 2025
The need for advanced information management technologies during pandemics is evident due to the rise in infectious diseases. This chapter explores natural language processing (NLP) techniques to categorize COVID-19-related tweets from X by five U.S. health organizations. A dataset of 2,500 tweets was manually annotated as COVID-related or not. Using Latent Dirichlet Allocation (LDA), COVID-related tweets (1,138) were further divided into five subcategories: Outbreak, Consequences, Prevention, Travel Advisory, and Resources. Multiple machine learning and deep learning models, including a BERT-based classifier, were employed for binary classification (COVID-related vs. non-COVID-related) and subcategory classification. The BERT classifier achieved the best F1-scores of 0.94 for binary classification and 0.88 for subcategory classification using individual label classifiers. While classifiers performed well for binary categorization, subcategory classification showed reduced performance due to imbalanced class labels. The use of LDA for topic modelling removes subjective bias in selecting topics while automatic classification of COVID-related tweets and sub-topics is essential for infodemic management. The findings emphasize the importance of advanced classifiers like BERT for accurate tweet categorization, aiding timely dissemination of crucial information. The proposed methodology provides a scalable framework for future outbreak preparedness and information management.