Addressing Imbalanced Data in Thyroid Disorder Diagnosis: A Novel Balancing Method and Hybrid Classifier

Pothuraju Raju, Bellamgubba Anoch, G. S. Naveen Kumar, Ramesh Babu Mallela, Thammuluri Rajesh, Ongole Pranathi Sudha · 2025

It might be hard to diagnose thyroid problems like hyperthyroidism and hypothyroidism because the data isn't always consistent and the symptoms can change. When working with datasets that aren't evenly distributed, traditional machine learning algorithms often lose their accuracy. This research suggests a Dynamic Selection Hybrid Model that works with a new BOOST Balancing Method to make diagnoses more accurate. Based on permutation feature importance, the model dynamically chooses the most relevant classifiers, such as KNN, SVM, Random Forest, AdaBoost, and Gradient Boosting, and then combines them using ensemble methods. To fix class imbalance, the BOOST approach combines SMOTE, Tomek Links, and sample weighting. Tests on a multi-class thyroid dataset with eight types of disorders show that the suggested model can reach an accuracy of 99.6%. The framework is more accurate and reliable than static models that have been used in the past. This approach is a flexible and scalable way to help doctors make decisions about thyroid treatment.

Read the paper · More papers on PaperTik