Web Phishing Detection Using Decision Tree Random Forest and XGBoost
Keerthana BP, Aparna Siva, T Senthilkumar, Thangavel Palanisamy, N. Prabhu, Kartik Srinivasan · 2025
The rise in web phishing attacks poses a critical threat to cybersecurity because they deceive users into disclosing personal information through fraudulent websites. This study compares three machine learning models for the identification of phishing websites based on important URL-based variables such as domain age, HTTPS presence, and lexical analysis: The Decision Tree provides an interpretable hierarchical classification, while the Random Forest builds resilience through multiple trees to improve accuracy and minimize overfitting. The gradient boosting algorithm XG Boost achieves substantial classification performance gains through boosting approaches [1]. The effectiveness of each model is evaluated by applying the ROC AUC analysis along with the measurements of the F1 score for accuracy, precision, and recall. ROC-AUC curves demonstrate how well each model can distinguish between phishing sites and legitimate websites. Experimental research shows that ensemble models like RF and XG Boost surpass individual decision trees in predictive capacity, which provides crucial insights for improving automated phishing detection systems. This research supports the development of reliable web security systems while highlighting machine learning as a key element in cybersecurity.