An Ensemble Learning Approach for Detecting Phishing Websites Using an Entropy-Based Feature Selection Method
Deepak P. Gupta, Ekta Gandotra, Meghna Dhalaria, Nivedita Gupta · Journal of Information & Knowledge Management · 2025
Phishing attacks have become increasingly common due to the growing number of web applications used in our day-to-day lives. The conventional tools and techniques for phishing detection are based on signature matching methods which do not have the capability to identify newly created phishing webpages. Thus, machine learning models are being used for this purpose. The performance of these models can be enhanced by considering a diverse and large number of features. However, building a machine learning model using high-dimensional data poses problems like increasing model building time and model overfitting. It hampers the accurate and timely detection of phishing webpages. Thus, it is important to employ the feature selection methods to shortlist the features without losing important information. It helps to develop the models with high performance in less time. In this study, the authors use information gain as feature selection method as part of data pre-processing. Six different models, i.e. Naïve Bayes, logistic regression, decision table, sequential minimal optimisation, [Formula: see text]-nearest neighbours, and random forest, are first built using the original set of 87 features and subsequently, with 30 features selected using the entropy-based information gain method. A comparative analysis is conducted on the basis of their efficiency and performance. Further, an ensemble model is proposed which combines the top three base classifiers built using the selected features. The experiments are conducted on a benchmark dataset of 11,430 instances containing 5,715 legitimate and 5,715 phishing websites to evaluate the effectiveness of the proposed model. The results reveal that the feature selection method improves the model building time of machine learning algorithms for detecting phishing websites without sacrificing accuracy. Further, the proposed stacked ensemble approach achieves an accuracy of 96.597%.