TubiLearn: Predictive Analysis of Tuberculosis Using Machine Learning Algorithms
Aman Das, Swaraj Bhattacharya, Lisa Roy, Liza Roy, Subhram Das, Papri Ghosh, Md Ashifuddin Mondal · 2025
Tuberculosis(TB) remains a critical global health challenge, particularly in resource-limited regions where access to timely and accurate diagnostic tools is constrained. This study introduces TubiLearn, a machine learning-based diagnostic framework designed to predict TB presence using synthetic clinical data. By leveraging advanced algorithms such as Random Forest, K-Nearest Neighbours (KNN), and Extreme Gradient Boosting (XGBoost), TubiLearn aims to enhance diagnostic accuracy and support clinical decision-making. A synthetic dataset of 10,000 samples was generated, incorporating key clinical features including age, gender, X-ray intensity scores, and symptoms such as cough duration, fever, and weight loss. To address class imbalance, preprocessing techniques like standardization, SMOTE, and downsampling were applied. The performance of the classifiers was evaluated using accuracy, precision, recall, F1 score, and the area under the receiver operating characteristic curve (AUC). Feature importance analysis identified the most influential predictors of TB. XGBoost demonstrated superior performance, achieving an accuracy of 89.6TubiLearn presents a scalable and cost-effective framework with significant potential for deployment in low-resource settings. By integrating synthetic data and state-of-the-art machine learning techniques, this framework provides a supplementary diagnostic tool that can assist clinicians in early detection and management of TB. Future work will focus on validating the framework with real-world datasets and exploring its integration into web-based diagnostic systems to further enhance its clinical utility.