Evaluating Causal AI Techniques for Health Misinformation Detection

Ommo Clark, Karuna Pande Joshi · 2025

The proliferation of health misinformation on social media, particularly regarding chronic conditions such as diabetes, hypertension, and obesity, poses significant public health risks. This study evaluates the feasibility of leveraging Natural Language Processing (NLP) techniques for real-time misinformation detection and classification, focusing on Reddit discussions. Using logistic regression as a baseline model, supplemented by Latent Dirichlet Allocation (LDA) for topic modeling and K-Means clustering, we identify clusters prone to misinformation. While the model achieved a 73% accuracy rate, its recall for misinformation was limited to 12%, reflecting challenges such as class imbalance and linguistic nuances. The findings underscore the importance of advanced NLP models, such as transformer-based architectures like BERT, and propose the integration of causal reasoning to enhance the interpretability and robustness of AI systems for public health interventions.

Read the paper · More papers on PaperTik