Data augmentation of JavaScript dataset using DCGAN and random seed
Ngoc Minh Phung, Mamoru Mimura · 2021
Many of the malicious JavaScript detection techniques focus on training a machine learning model with balanced dataset. However, real-world data only has a small fraction of malicious JavaScript, making it an imbalanced dataset. Training and testing the model with an imbalanced dataset can help evaluate the practical uses of that model. This paper aims to build a filter model that can quickly classify JavaScript malware using natural language processing (NLP) and machine learning. The feature of words inside the JavaScript file will be converted into vector form and used to train the SVM classifier. Different NLP models and oversampling methods are tested to achieve high recall score, such as Doc2Vec and Latent semantic indexing (LSI) model. In this paper, DCGAN model will be used to generate new training malicious data based on original training dataset. The goal is to let the DCGAN model learn the feature of the training malicious data and generate more reliable data than previous experiment which use random seed for oversampling. We evaluate our models with a dataset of over 20,000 samples obtained from top popular websites, PhishTank and other source. The experimental result shows that the best recall score achieves 0.78 with the LSI model.