Explainable AI for Public Health: Predictive Modelling of Non-Communicable Diseases Risk using SHAP and LIME

Md Sarfaraz Ahmad, R. Aroul Canessane · 2025

Non-communicable diseases (NCDs) are a grave threat to public health in India, as they are a prime contributor to overall morbidity and mortality rates in the country. The economic impact of NCDs in India also poses a lot of challenges, sapping precious and limited resources the country has at its disposal. As the country finds itself in a precarious situation with ever-increasing cases of dangerous diseases like diabetes, cardiovascular problems, and cancer, the need for a data-driven approach to detect and mitigate the impact of these dreadful ailments in the very initial stages assumes an added urgency. Artificial Intelligence (AI) and Machine Learning (ML) have given cause for a lot of hope, demonstrating immense potential for scalable, population-wide predictive modeling. However, despite the promise and the ability of AI and ML, their clinical deployment falls short of expectations. The reasons cited for these include the high level of opacity surrounding many high-performing models, which often leads to a lack of trust, doubtful interpretability, and inability to adhere completely to regulatory requirements. This study deals with a vigorous and explainable predictive framework that integrates six widely used ML classifiers Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Random Forest (RF), Multi-Layer Perceptron (MLP), XGBoost, and Deep Neural Networks (DNN)—trained on comprehensive Indian health datasets including NFHS-5, National Health Profile (NHP), and ICMR reports. Feature selection was optimized through techniques such as Elastic Net, SVR-RFE, and SBFS-RF. To ensure the desired transparency, we employed state-of-the-art explainable AI (XAI) techniques SHAP and LIME for both global and local interpretability. Our experiments demonstrated that DNN and XGBoost both constantly achieved top-tier accuracy (95.01% and 94.11%, respectively). However, the higher level of complexity associated with them often inhibited transparency. On the other hand, RF seamlessly combines strong performance (90.01%) with interpretability, and that too with a lower requirement for a higher degree of computational power. This makes it a better candidate for deployment in real-world clinical environments. This work underlines the all-important role of balancing accuracy, credibility, and interpretability in the realm of AI-driven healthcare, while presenting detailed and practical insights for the realistic integration of ML solutions into India’s overburdened and stretched public health infrastructure

Read the paper · More papers on PaperTik