Identification and Classification of Malicious and Benign URL using Machine Learning Classifiers
Danish Kundra · 2023
In a world where digitalization is constantly advancing and developing, cybersecurity and cyberwarfare are essential. Malware has grown to be a serious menace to internet users in the present digital age. Malware has a rapid rate of propagation and is a serious danger to online safety. Network security measures are therefore crucial for thwarting these online dangers. With a Universal Resource Locator(URL), communities can access the World Wide Web resources that are vital to daily lives). Attackers use such channels of communication and malicious URLs to carry out scams and trick people by setting up misleading and false websites and domains. These dangers let in a wide range of dangerous attacks, including spam, malware, phishing, and spyware. In order to stop the emergence of numerous cybercriminal acts, it is vital to identify dangerous URLs. In this study, collection of machine learning (ML) algorithms are analyzed to identify harmful websites using a dataset made up of 64,1119 URL records. When compared to other classifiers, the random forest classifier performs at a 91.49% accuracy rate better than gradient boost, extreme gradient boost, adaptive boost, and k-nearest neighbor machine learning techniques. This research has contributed to the area by carrying out extensive feature extraction and analysis to determine the best characteristics for hazardous URL predictions, contrast several models, and reach a high level of accuracy by utilizing a sizable new URL dataset.