Malicious URL Analysis Using a Suite of Machine Learning Techniques
M. Balaji, R.Jospehine Sahana, Rajkumar N · 2023
Recognizing and classifying dangerous URLs is essential for assuring user safety in light of the exponential growth of Internet usage. Numerous cyber threats, such as phishing, malware distribution, and identity theft, are enabled by malicious URLs. This research seeks to enhance user safety by developing a classification algorithm based on machine learning that can differentiate between malicious and secure URLs. A massive database of potentially dangerous and secure URLs is analyzed in the first stage of the research process. The URLs are preprocessed to eliminate redundant information and extract pertinent properties (such as length, symbols, numerals, and characters). Frequently, the dataset is divided into training and test sets so that various classifiers can be compared directly. This study uses six distinct machine learning algorithms to detect and classify malicious URLs: Random Forest, XGBoost, Support vector machine, convolutional neural network, recurrent neural network, and artificial neural network. The results that cast light on the algorithm's performance are accuracy, precision, recall, and F1 scores. The study discovered that the recommended CNN model distinguished malicious from secure URLs with a success rate of 99.9%. Examining criteria such as URL length and symbol usage significantly improves classification accuracy. The research also examines whether URL IP addresses can be used as indicators of malicious intent. This research's findings can be applied to detecting and preventing intrusions employing malicious URLs. This study arrived at the chosen categorization approach through comparison; it will serve as the basis for future research and real-world implementations in cybersecurity. Increases user security against cyberattacks and data intrusions by precisely labeling and detecting problematic URLs.