Hybrid Approach for Malicious URL Detection by Integrating TF-IDF and Random Forest Techniques

J. Benita, Chavidi Balaji, Damarouthu Kamalesh, Darvemula Sreeram Krishna, Chunduri Mohan Narasimharao · 2024

The Internet has become an essential component of daily life due to its broad use, but it has also rendered websites more vulnerable to cyberattacks. Current security measures frequently come up lacking in detecting and defeating these threats, underscoring the need for quicker and more precise detection techniques. This paper suggests a machine learning-based approach for URL analysis-based dangerous website detection. This study suggests a hybrid machine learning-based method for categorizing URLs into four groups: malware, phishing, defacement, and benign. Using a dataset that has been processed by a TF-IDF vectorizer, a Random Forest Classifier is trained. The model and vectorizer that are produced are then saved using Joblib for effective reuse. TF-IDF vectorizes the user-provided URLs, and the Random Forest Classifier is used to determine the outcome of the trained model and Vectorizer. This suggested method's high reliability was demonstrated by its 96% accuracy in categorizing URLs into benign, defacement, phishing, and malware groups. Interestingly, only 10% of the dataset was used to train the model, demonstrating its scalability and effectiveness. This study highlights the possibility of using Random Forest classification in conjunction with TF-IDF vectorization as a simple yet powerful way to identify harmful websites. This approach greatly improves web security by offering a strong tool for early threat identification and mitigation due to its high accuracy and low data requirements, making it appropriate for real-time applications.

Read the paper · More papers on PaperTik