Comparing Different Evaluation Metrics with the Grid Search Method for Classification of Highly Imbalanced Data

Tsvetelina Kaneva, Boris Ivanov Evstatiev, Irena Valova, Nikolay Valov, Katerina Gabrovska-Evstatieva · 2024

In machine learning, imbalanced data is a frequent challenge, often leading to models that excel in predicting maj ority classes but perform poorly on minority classes. This study examines the effectiveness of different evaluation metrics and the use of the Grid Search method for hyperparameter tuning when handling imbalanced datasets in multiclass classification problems. Utilizing data from cattle behavior classification, several machine learning algorithms—Random Forest Ensemble (RFE), Decision Tree (DT), Support Vector Machines (SVM), and Naïve Bayes (NB)—are evaluated. The study reveals that commonly used metrics, like accuracy, can be misleading in imbalanced scenarios, suggesting the need for alternative metrics such as balanced accuracy, precision, recall, and Fl score. Results demonstrate that RFE and DT are more robust and reliable in handling class imbalance, whereas SVM and NB models show significant limitations. This paper underscores the importance of carefully selecting evaluation metrics and optimizing hyperparameters to achieve dependable outcomes in machine learning applications involving imbalanced data.

Read the paper · More papers on PaperTik