Performance of RUS and SMOTE Method on Twitter Spam Data Using Random Forest

Huda Ubaya, Ria Siti Juairiah · Journal of Physics Conference Series · 2020

Abstract Spam data on Twitter has one problem, which is imbalanced data. This study proposes two resampling methods, namely RUS and SMOTE as one way to balance the data. In RUS, the data becomes less because random data is deleted while the data on SMOTE will be even greater because synthetic data is grown. To measure the impact of these two methods on the dataset, the Random Forest Classifier method will be applied. In the results, RUS and SMOTE can increase the value of the Area under Curve to 95.21% and 99.13%. The best performance is obtained when applying a balanced ratio of SMOTE.

Read the paper · More papers on PaperTik