Handling of unbalanced LC-MS medicinal plant data using Near-Miss Undersampling tested with Gaussian Naive Bayes and K-Nearest Neighbors
Iwan Binanto, Rosalia Arum Kumalasanti, Nesti Fronika Sianipar · 2023
Imbalanced data refers to data with classes that have extreme majority and minority data. Such data can lead to inaccurate results. The dataset used in this study came from LC-MS data of medicinal plants that had previously been labeled using the webscraping method and unbalance. There are several resampling algorithms to balance the data. This study used nearmiss undersampling with consideration for being more robust against overfitting. The balanced data was split for training and testing with a ratio of 70:30, which will be tested using Gaussian Nave Bayes and K-Nearest Neighbors classification algorithms. The results showed that Near Miss version 1 sampling with the Gaussian Naive Bayes algorithm provided better accuracy and faster execution time.