Machine Learning Algorithms for Phishing Detection: A Comparative Analysis of SVM, Random Forest, and CatBoost Models

Preet Deep Singh, Taniya Hasija, KR Ramkumar · 2024

Phishing attempts are increasing nowadays due to the advanced internet penetration worldwide. The objective of this research is to create a phishing detection system that is both efficient and effective. This work has taken greater significance in enhancing cybersecurity against the increasing threat of phishing attacks. The proposed approach is for data preparation and applying SMOTE for dataset balancing, training/ testing on a dataset of 11,430 URLs with 87 features extracted from URL structure, content, and external services. The findings indicate that CatBoost gives the best accuracy, thus performing better with fewer misclassifications than Random Forest, evidenced by its confusion matrix. The research concludes that integration herewith of machine learning models more so CatBoost results in phishing detection systems has improved its accuracy and reliability to an accuracy of 97.24%. Such advancement provides a robust solution for mitigating phishing risks.

Read the paper · More papers on PaperTik