A URL address aware classification of malicious websites for online security during web-surfing
Goutam Chakraborty, Tsai Tzung Lin · 2017
On the Internet, users often visit unknown websites. However, malicious websites are a significant threat to the Internet users. Malicious websites implant malwares into users computers without their knowledge, through drive-by-downloads technology. A naive user could easily fall victim of such attack. With increased use of internet browsing, web security is an important issue and an important research topic. The motivation of this study is to classify malicious web-sites from benign ones from their URL features. If it could be done with high precision, especially with low false accept rate, automatic blocking of suspicious URL at the user site will be possible. We collect URL data for a large number of known benign as well as malicious websites. URL data has many characteristics. Some of them are relevant to classify the site as malicious and others are not. Many characteristic features are textual. We first converted all such features into suitable numeric data relevant to that feature, so that the numeric values truly represent the information of the original feature. We then select the relevant features for our classification task, by using two methods: (1) least absolute shrinkage and selection operator (LASSO) and (2) Multi-objective Pareto Genetic algorithm (MOGA). Finally, the data consisting of selected features is used to train a support vector machine (SVM) classifier. A ten-fold validation is used to estimate the performance. Performances of two feature selection methods as well as another recently published report were compared. By feature selection using Pareto GA, we could achieve more than 95% classification accuracy and F-score with least number of features.