Interpretable Illness-Category Classification from Drug Attributes Using XGBoost with SHAP Explanations: A Study on the Pharma-Safe Index Dataset
Bushra Hoque, Rafi Uddin, Karib Shams, Md Miskat Hossain, Farhana Ahmed Tasnim, Mohammad Rifat Ahmmad Rashid, Raihan Ul Islam, M. Saddam Hossain Khan · 2025
Accurately classifying illness categories based solely on drug-related information presents significant challenges, as inferring the associated illness from drug data alone can often be complex and uncertain. To address this, our research utilized various machine learning models to classify illness categories based on drug-related features, incorporating explainable AI (XAI) techniques to enhance interpretability. Using the Pharmasafe Index Dataset, which contains pharmaceutical data from Indonesia, we trained and evaluated several machine-learning models, including Support Vector Machine (SVM), Logistic Regression, K-Nearest Neighbors (KNN), Neural Networks, XGBoost, and Random Forest. Our experimental results show that XGBoost achieved the highest classification accuracy of 99.52 %, while SVM demonstrated the lowest accuracy at 73.16%. To further interpret the predictions of the top-performing model, we applied SHAP (SHapley Additive exPlanations) values, providing insights into how the model makes classification decisions. Additionally, chi-square testing was employed to identify significant relationships between various features in the dataset. Through this framework, users can more easily determine potential illness categories based on the drugs they are prescribed, improving diagnostic support and decision-making.