A Comparative Insight into Machine Learning Strategies for Phishing Website Identification
Preet Deep Singh, Taniya Hasija, KR Ramkumar · 2024
Phishing attacks have posed a serious threat to the information security cyber world; where malicious websites try to beguile the actual identity of users. Attackers keep adopting new tactics for phishing sites over a while, making it challenging to detect new phishing sites. It has been based on a comparative study of various machine learning models, such as SVM, RandomForest (RF), CatBoost, Voting Classifier, and Stacking Classifier, toward their classification of phishing and non-phishing sites. Different experiments were performed on the dataset, which consisted of features of URLs. Accuracy, precision, recall, and F1-score measure performance. Among the various models used for evaluation, the best-performing classifier was the Stacking Classifier with an accuracy of 97.33%, precision value of 0.9684, recall of 0.9793, and F1 score of 0.9738 for the non-phishing websites. This research study has outperformed the rest against phishing websites with a precision of 0.9785, recall of 0.9672, and F1 score of 0.9728. The ensemble approach due to the strengths of multiple base learners improves the classification performance much more than that of any of the individual base learners. This work realizes that there is a great need for future work based on the effectiveness of stacking techniques in phishing detection, and it has the potential to open the door for methods to enhance real-time detection systems against adversarial threats.