A Machine Learning approach for Detecting Malicious URL using different algorithms and NLP techniques

M.A. Waheed, Baswaraj Gadgay, D C Shubhangi, P Vishwanath, Qurat Ul Ain · 2022

A malicious URL is one that was made specifically to attack through spam or fraud. Due to the billions of dollars that are compromised, malicious URLs pose a severe threat to security software. Finding secure and phishing links is therefore crucial. Therefore, machine learning is quite helpful for resolving security-related challenges. In this study, we use about 5 lakh URLs that were retrieved from the Kaggle dataset. We are utilizing three NLP approaches, including the count vectorizer, hash vectorizer, and TF-IDF vectorizer. Six machine learning classifiers, including the decision tree, random forest, K-NN, NB, SVM, and logistic regression, were used in conjunction with all these techniques. The highest accuracy results of 98.2 percent are produced by random forest. To determine whether or not the URL supplied is malicious, we built a web app using Flask.

Read the paper · More papers on PaperTik