Phishing Web Page Detection using Web Scraping

Mallika Boyapati, Ramazan Savas Aygün · 2023

In today’s Internet era, webpages act as major user interfaces for the applications hosted by the organizations. One of the most common security attacks on the web page applications is phishing attacks. Phishing is a social engineering attack in which an adversary tries to steal user credentials by tricking them to believe that they are on a legitimate web page. Adversaries are using sophisticated and new ways to forge the web page designs craftily to trick the users into visiting the malicious links. The phishing webpages are used as a medium to carry out the art of phishing attacks. Web scraping is a methodology to extract features of each webpage. In this paper, web scraping is employed to extract hybrid feature set to implement Machine Learning (ML) models. The machine learning models like XGBoost, Multilayer Perceptron, Logistic Regression, SVM, Auto Encoder, Random Forest, Decision Trees, K-means, and Naive Bayes are built. All the models are tested with and without applying principal component analysis (PCA), a feature reduction technique. Extracting hybrid features by employing web scraping to train the ML models and finding the key features contributing to phishing detection on Phishpedia, Kaggle, and PhishTank datasets is the key contribution of this paper. Results show that XGBoost algorithm outperformed the all other classifiers with 98% accuracy or higher on the web scraped features on all the datasets. The model achieved higher precision and recall when compared to other approaches like CANTINA+, URL based approaches, SenseInput, Knowing Thy Domain, and Phishpedia on the Phishpedia dataset.

Read the paper · More papers on PaperTik