Explainable Machine Learning for Phishing Site Detection: A High‐Efficiency Approach Using Boosting Models and SHAP

Khandaker Mohammad Mohi Uddin, Nitish Biswas, Sarreha Tasmin Rikta, Md. Nur ‐A‐Alam, Rafid Mostafiz · The Journal of Engineering · 2025

ABSTRACT Nowadays, a wide range of electronic devices are being connected through the internet, and a wide range of cybercrimes are being coordinated. Phishing is one of the most serious online crimes. It poses a major threat to people and businesses, resulting in enormous financial losses and user‐provided data breaches. Traditional phishing detection methods are time‐consuming and often fail to detect novel phishing approaches. This study aims to develop an efficient and explainable machine learning model to detect phishing websites, providing transparency in predictions to increase user trust. To tackle this challenge, we suggest a highly efficient phishing detection method that uses ensemble learning techniques, specifically gradient boosting machine, extreme gradient boosting, and light gradient boosting, combined with Shapley additive explanations for clarity. The innovation comes from integrating SHAP with ensemble models to provide high accuracy and clear explanations, which helps end‐users and security experts comprehend the decision‐making process. Using a publicly available phishing dataset from Kaggle, the XGBoost model achieved a high accuracy of 97.78%. It outperformed other models. The SHAP‐based analysis visually explained the features of the importance and contribution of each classifier. This approach promotes model transparency and improves the reliability of the phishing detection system.

Read the paper · More papers on PaperTik