A Comparison of Random Forest and Support Vector Machine Classification Algorithms for Imbalanced and Balanced Rodent Tuber Dataset with Random Oversampling Method
Iwan Binanto, Marselinus Sandimus Jamlu, Robertus Denyva Adibuana Wibisono, Nesti Fronika Sianipar · 2024
Class distributions that are assumed to be balanced by conventional machine learning algorithms make them susceptible to biases that favor the dominant class throughout the classification process. This study primarily discusses how classification results may be impacted by the dataset’s state, whether it is balanced or unbalanced. The Rodent Tuber’s unbalanced dataset, which includes chemical data useful for cancer diagnosis, was utilized in this investigation. Support Vector Machine with three kernels and Random Forest will be used to classify this dataset. To see the differences, Random Oversampling is used to do classification on both an imbalanced and balanced dataset. The outcomes demonstrate that better results with shorter running times are obtained from a balanced dataset. The Random Forest method yielded the best classification results on the balanced dataset.