PhishHunter-XLD: An ensemble approach integrating machine learning and deep learning for phishing URL classification

Tirth Doshi, Vishva Patel, Nemil Shah, Debabrata Swain, Debabrata Swain, Debabala Swain, Debabala Swain, Biswaranjan Acharya · Franklin Open · 2025

Phishing continues to pose a significant cybersecurity threat by deceiving users into disclosing sensitive information through maliciously crafted URLs. Traditional detection methods, including blacklists and heuristic analyses, have proven inadequate against evolving phishing techniques due to their reliance on static patterns and manual updates. In this study, a weighted voting ensemble framework has been proposed, integrating semantic feature extraction using DistilBERT with classical machine learning classifiers (XGBoost) and deep learning models (LSTM) to enhance phishing URL detection. Model complementarity has been leveraged XGBoost captures explicit lexical features, LSTM models sequential dependencies, and DistilBERT extracts contextual semantics resulting in an adaptive decision boundary that improves generalization and reduces false positives. Extensive experiments conducted on large-scale benchmark datasets, such as the “Phishing Site URLs” and “Malicious URLs” datasets, have demonstrated that the proposed ensemble framework achieves a detection accuracy of 99.83% with low computational latency. Furthermore, the system has been deployed via Streamlit, providing a real time, interactive interface for cybersecurity practitioners. Future work will explore optimization strategies, including model pruning, quantization, and adversarial training, to further enhance efficiency, scalability, and resilience against emerging zero-day phishing techniques.

Read the paper · More papers on PaperTik