PhishRescue: A Stacked Ensemble Model to Identify Phishing Website Using Lexical Features
Fahima Sharmin Hossain, Linta Islam, Mohammed Nasir Uddin · 2022
The increased usage of the internet due to the development of network technologies has increased the probability of social engineering attacks in different sectors such as financial institutions, webmail, payment, social media, retail, shipping, and cryptocurrency. Out of five common social engineering attacks such as phishing, pretexting, baiting, quid pro quo, and tailgating, phishing is the most familiar one. Massive loss of money has occurred because of this attack. So, several approaches have been proposed till now to fight this attack. Among all the approaches, a machine learning algorithm is an appropriate method to identify phishing attacks because of its ability to find patterns in URL (Uniform Resource Locator) to detect phishing sites. This study uses two datasets of URLs from Kaggle. During preprocessing, we removed unnecessary values and then extracted the appropriate attributes to accurately classify the phishing URL from all types of URLs using regular expressions in natural language techniques. Finally, we used six individual classification algorithms: Extra Tree, Logistic Regression, Gaussian Naïve Bayes, Decision Tree, Random Forest, Gradient Boosting, and three ensemble methods such as Bagging Stacking and Voting for dataset training and testing. The results of our tests show that the proposed model achieves an accuracy of 85.08% for binary classification and 81.31% for multi-class classification for the first dataset and an accuracy of 95.99% for binary classification for the second dataset