Tuberculosis prediction: performance analysis of machine learning models for early diagnosis and screening using symptom severity level data
Suresh S, Dhanalakshmi S · International Journal of Basic and Applied Sciences · 2025
Tuberculosis (TB) remains a formidable issue for worldwide public health and calls for swift and exact diagnostic strategies to achieve the best health results for those affected. A methodical machine learning (ML) sequence was diligently followed, featuring data preprocessing, feature choice, encoding, and the training of the model in a logical order. A detailed investigation was performed on six unique machine learning architectures, comprising the ANN, SVM, Decision Tree, Random Forest, XGBoost, and Logistic Regression, closely analyzing their key performance measures essential for measuring their effectiveness, including accuracy, precision, recall, F1-score, and AUC-ROC, hence providing an extensive view of their attributes and feasible uses across different sectors. The matter of class imbalance was diligently approached through the execution of the Synthetic Minority Over-sampling Technique (SMOTE), and the model's performance was scruti-nized using 5-Fold Cross-Validation to affirm both consistency and relevance of the conclusions. Achieving a stellar accuracy of 99.55%, an impeccable recall of 100%, and a noteworthy F1-score of 99.54%, the ANN model is hailed as the premier model for tuberculosis forecasting. The Random Forest and SVM models also illustrated robust predictive performance, evidenced by elevated accuracy and AUC-ROC scores. In a contrasting view, Logistic Regression provided the least successful outcomes, suggesting that linear models could be inadequately matched to the attributes of this dataset. This study elucidates the efficacy of machine learning methodologies in the diagnostics of TB and emphasizes the critical role of symptom analysis and data-informed decision-making within the healthcare sector.