Malicious URL Detection: Comparative Study of Machine Learning Algorithms

Thadakaluru Jaswanthi, Talluri Harshitha, Simhadri Tanya, Meena Belwal · 2024

Malicious URLs can be detected by leveraging lexical analysis and machine learning, where the parsed components of the URL are parsed, and its features include domain reputation, URL length, and the presence of suspicious characters. Research on the detection of malicious URLs suffers challenges in keeping pace with the changing structural landscape of URLs. Adaptation to dynamic patterns and addressing the technical issues of efficient processing and scalability become important to stay one step ahead of the evolving threat landscape. This paper concentrates on addressing these challenges with the integration of Machine Learning and Lexical Analysis. Features extracted in this process are then fed into machine learning algorithms for the classification of URLs and are tested on unseen data. Research performance is measured using several metrics including Precision, Recall, F1-score and Accuracy during the Model evaluation. Out of all the models experimented, Extra Trees classifier stands out the best with 92% accuracy, showing effectiveness in accurate detection of malicious URL’s.

Read the paper · More papers on PaperTik