“Boosting Tuberculosis Classification Accuracy with Polynomial Features and Random Forests”
Parvez Rahi, Inderjeet Singh, Anushka, Ankit Kumar, Tarun Baliyan · 2024
From the WHO data of the current year, one realizes that tuberculosis is still a world- wide health issue with about 10 million new cases and over 1.5 million deaths annually. TB, in its various forms, comprises latent tuberculosis, extra pulmonary tuberculosis (EPTB), miliary tuberculosis and pulmonary tuberculosis (PTB), which constitutes 85% of cases. It is crucial to delineate these types correctly in order to decrease mortality and enhance the quality of treatments. This work modifies the Random Forest classifier to enhance the TB classification model's accuracy. To deal with the imbalanced data, Synthetic Minority Oversampling (SMOTE) was used while Polynomial Feature Engineering ensured sufficient feature interactions in the data set. The preprocessing of dataset included feature scaling and label encoding while a major effort was made towards hyperparameter optimization using GridSearchCV. Our model attained a high test accuracy of 92% having MSE of 0. 08. Some of the evaluation criteria used to inspect performance highlight on how the model differentiates between TB types and it achieves a precision of 0.91, recall of 0.64, and F1 score of 0 the algorithm achieved best performance on files with less than 92.91. Subsequent studies could employ more clinical information about the patient in order to improve the model and its generalization across different people of different ages, genders or of different origins.