Implementasi SMOTE dan Support Vector Machine Pada Klasifikasi Data Tidak Seimbang Metilasi Arginin

Favorisen Rosyking Lumbanraja, Ester Caroline Lumban Gaol, Dewi Asiah Shofiana, Akmal Junaidi · Jurnal Pepadun · 2024

Imbalanced data is one of the crucial problems in machine learning and data mining which may provide low accuracy in minority classes and makes the classification method not fully optimized. The Arginine Methylation dataset for example, gives a large amount of imbalanced data. Methylation is one of the post-translational modification processes that occurs in arginine protein which affects signal transduction and RNA binding inside cytoplasms. Therefore, it is essential to handle imbalanced data for classification. Synthetic Minority Oversampling Technique (SMOTE) is an algorithm for solving imbalanced data in classification using the concept of k-nearest neighbors. Support Vector Machine (SVM) is a supervised learning method which splits datasets using hyperplane and maximize margin distance. In this research, the arginine methylation dataset is divided into three experimental data, which consists of training data, testing data, and independen data. Data processing goes through a series of steps; data pre-processing (clean redundance data), feature extraction (generates 159 feature dimensions), SMOTE and SVM modeling, and classification testing using 10-fold cross-validation and confusion matrix. The accuracy of training data is 100% in RBF kernel, whereas testing data gives a low accuracy of 65,90% in linear kernel. Independen data have decent accuracy in linear kernel by 98,50% percentage.

Read the paper · More papers on PaperTik