A Robust Voting-ML System in Identifying Phishing Websites Using Hybrid Ensemble Learning and Advanced Feature Selection Techniques
Amal El Arid, Ghalia Nassreddine · 2025
Phishing remains a prevalent cyber threat where attackers impersonate legitimate organizations to deceive individuals into disclosing sensitive information via communication platforms. Even with sophisticated security systems, phishing exploits human vulnerabilities, leading to potential identity theft, unauthorized transactions, and significant financial and reputational harm. Advanced phishing detection systems employ techniques such as behavioral analysis, user profiles, and anomaly detection to improve the accuracy of the identification. This paper describes a strong model for finding phishing websites. It uses supervised machine learning algorithms like adaptive boost, gradient-boosted trees, k-nearest neighbors, Gaussian naive Bayes, random forest, and extreme gradient boost. The model was trained on a dataset of$\mathbf{5, 0 0 0}$real websites and$\mathbf{5, 0 0 0}$fake websites with 48 extracted features. This method is unique because it uses a new feature selection strategy that combines the results of three different feature engineering techniques: mutual information regression, principal component analysis, and truncated singular value decomposition. These selected features are then scaled using the MinMaxScaler function. Additionally, a majority-voting ensemble is applied to integrate the predictions from the various machine-learning classifiers, thereby enhancing overall detection performance. The proposed system's effectiveness is evaluated using accuracy, precision, recall, and f1-score, demonstrating its capability in distinguishing phishing websites.