Optimization of Phishing Website Classification Based on Synthetic Minority Oversampling Technique and Feature Selection
Rizal Dwi Prayogo, Siti Amatullah Karimah · 2020
This paper presents a new approach for optimizing phishing website classification based on Synthetic Minority Oversampling Technique (SMOTE) together with feature selection. Classification is a kind of supervised machine learning technique that learns based on the features to identify the class. However, not all features are relevant to identify phishing websites and the class imbalance problem leads to suboptimal performances. Therefore, we propose SMOTE for handling the class imbalance problem by generating new synthetic instances for the minority class. Filter-based feature selection using Information Gain and Correlation are proposed for reducing irrelevant features. The classification performances are evaluated using K-Nearest Neighbor (KNN) classifier. The results demonstrate that SMOTE effectively increases the classification performances in terms of accuracy, precision, recall, and F-measure with more time-efficient. The performance of SMOTE combined with feature selection is validated and benchmarked with different techniques both on full features and reduced features. The results demonstrate that our proposed technique presents the highest accuracy, i.e. 97.47% on full features and 94.87% on reduced features. Hence, our proposed technique is promising in optimizing phishing website classification.