MalDeWe: New Malware Website Detector Model based on Natural Language Processing using Balanced Dataset
Jovana Dobreva, Aleksandra Popovska‐Mitrovikj, Vesna Dimitrova · 2021 International Conference on Computational Science and Computational Intelligence (CSCI) · 2021
The increasing use of the Internet in everyday life and the huge number of users leads to increasing the number of malicious websites, which aim is to damage a computer system or compromise data without the owner’s consent. In this paper we propose a Natural Language Processing (NLP) model, called Malware Website Detector (MalDeWe), for malware web page identification that is trained on a domain-specific corpus. The major goal of this model is to transfer English word knowledge from the pre-trained model RoBERTa into a dataset of JavaScript codes that included the website’s text context as well as certain JavaScript expressions. With this model, we obtain a Roc Auc score of 0.95. Therefore, we can conclude that our model is doing admirably in terms of identifying a malicious web page. On the other hand, we may infer that one of the most essential aspects to consider while training a classification model is dataset balance. Whereas the model trained on the initial unbalanced dataset failed to detect harmful websites, the model trained on the balanced dataset, correctly identified 95% of dangerous websites. In that sequence, we may deduce that the metrics used for the model evaluation are critical. Therefore, we recommend using the Roc Auc score, Recall, Precision, and Confusion matrix as evaluation metrics. So, in this paper we propose a new NLP model, and we discuss the reasons why the choice of evaluation metrics are important and how the dataset balance makes changes on the model efficacy.