Comparative Analysis and Practical Implementation of Machine Learning Algorithms for Phishing Website Detection
Samad Najjar-Ghabel, Shamim Yousefi, Payam Habibi · 2024
Phishing, characterized by fraudulent attempts to obtain sensitive information through deceptive communications, poses a significant cybersecurity threat. The rapid increase in internet usage and the proliferation of online services have made phishing attacks more prevalent and sophisticated. Machine learning algorithms have emerged as powerful tools in identifying and mitigating these threats by analyzing vast amounts of data and detecting patterns indicative of phishing activities. Various studies have examined and compared machine learning-based methods for phishing website detection, but they often lack practical implementation and do not comprehensively compare criteria and classification algorithms on specific datasets. This paper aims to enhance phishing website detection by assessing the performance of six distinct classification algorithms: k-Nearest Neighbors, Naive Bayes, Random Forest, Bernoulli Naive Bayes, Support Vector Machine, and Decision Tree. Utilizing the “Phishing Websites” dataset from the UCI Machine Learning Repository, the classifiers are evaluated based on F-Score, accuracy, precision, and recall. Among these algorithms, Random Forest demonstrates superior performance, achieving an accuracy of 96.7%, precision of 96.1%, recall of 98.2%, and an F-Score of 97.1%. These results underscore the robustness and efficacy of Random Forest in phishing website detection tasks. While other algorithms also show promise, especially under specific conditions, further research is needed to explore the potential improvements from combining multiple algorithms and examining additional features. This direction offers a promising avenue for enhancing phishing website detection systems in future studies.