Tuberculosis prediction: performance analysis of machine learning models for early diagnosis and screening using symptom severity level data

Suresh S, Dhanalakshmi S · International Journal of Basic and Applied Sciences · 2025

Tuberculosis (TB) remains a formidable issue for worldwide public health and calls for swift and exact diagnostic strategies to achieve the ‎best health results for those affected. A methodical machine learning (ML) sequence was diligently followed, featuring data preprocessing, ‎feature choice, encoding, and the training of the model in a logical order. A detailed investigation was performed on six unique machine ‎learning architectures, comprising the ANN, SVM, Decision Tree, Random Forest, XGBoost, and Logistic Regression, closely analyzing ‎their key performance measures essential for measuring their effectiveness, including accuracy, precision, recall, F1-score, and AUC-ROC, ‎hence providing an extensive view of their attributes and feasible uses across different sectors. The matter of class imbalance was diligently ‎approached through the execution of the Synthetic Minority Over-sampling Technique (SMOTE), and the model's performance was scruti-‎nized using 5-Fold Cross-Validation to affirm both consistency and relevance of the conclusions.‎ Achieving a stellar accuracy of 99.55%, an impeccable recall of 100%, and a noteworthy F1-score of 99.54%, the ANN model is hailed as ‎the premier model for tuberculosis forecasting. The Random Forest and SVM models also illustrated robust predictive performance, evidenced by elevated accuracy and AUC-ROC scores. In a contrasting view, Logistic Regression provided the least successful outcomes, ‎suggesting that linear models could be inadequately matched to the attributes of this dataset. This study elucidates the efficacy of machine ‎learning methodologies in the diagnostics of TB and emphasizes the critical role of symptom analysis and data-informed decision-making ‎within the healthcare sector‎.

Read the paper · More papers on PaperTik