Detection of Multi-Class Website URLs Using Machine Learning Algorithms
Pandu Ranga Raju B · International Journal of Advanced Trends in Computer Science and Engineering · 2020
Online phishing is one of the Internet's most widespread crime schemes.A common counter measure includes testing URLs against blacklists of established phishing websites, which are typically collected on the basis of manual verification and are inefficient.As the Internet scale expands, automatic URL detection is increasingly necessary to provide timely security for end-users.In this paper, we propose an efficient and versatile malicious URL detection system with a rich collection of features representing the diverse characteristics of phishing websites and their hosting platforms, including features that are difficult to forge.Using the Random Forests algorithm, our program benefits from both high detection capacity and low error levels.Based on our experience, this is the first research to carry out these large-scale websites / URL scanning and classification experiments, taking advantage of the distributed viewing points for the feature set.The results of the experiment show that our program can be used by the blacklist provider to create automated blacklists.