Improving the Performance of Semantic-Based Phishing Detection System Through Ensemble Learning Method
Alisha Maini, Navan Kakwani, Bandi Ranjitha, M. K. Shreya, R. K. Bharathi · 2021 IEEE Mysore Sub Section International Conference (MysuruCon) · 2021
Technology is evolving at an exponential rate, and so are human minds. The innovativeness in technology has both its advantages and disadvantages. It has made general work easier and comfortable by building cyberspace but it has also led to cybercrimes. One of the cybercrimes is phishing attacks. Traditional anti-phishing techniques which use blacklists to iterate and check if the URL is legitimate or phishing is not very useful as the phishers can attack using new URLs. Therefore, Machine learning algorithms can be used to train models to learn the semantic differences between legitimate and phishing URLs. To perform classification of legitimate and phishing URLs, eight ML algorithms which are Random Forest, Decision tree, Naive Bayes, AdaBoost, KNN, XGBoost, Support Vector Machines (SVM) and Logistic Regression are trained and tested. To improve the standard of the classification model, an ensemble model is built using the above-mentioned machine learning algorithms. From the results observed, amid the machine learning algorithms, XGBoost achieved the highest accuracy and the ensemble model achieved an accuracy higher than all individual machine learning models.