Evading VBA Malware Classification using Model Extraction Attacks and Stochastic Search Methods

Brian C. Fehrman, Francis E. Akowuah, Randy C. Hoover · 2024

Antivirus (AV) software that relies on learning-based methods is potentially vulnerable to adversarial attacks from threat actors. Threat actors can utilize model-extraction attacks against AV software to create a surrogate model. Malware samples can be tested against the surrogate model to determine how the target AV software will classify a given sample. Using a surrogate model speeds the process of malware development by allowing modifications to first be tested in feature space, which is significantly faster than performing modifications in code space. This work investigates performing evasion attacks against Windows Defender VBA malware classifier in an offline mode. The performance of five machine learning models is compared for their use as surrogate models. The models are reinforced by augmenting their training sets with samples that are generated by modifying existing samples. The results show that model performance is greatly improved with the augmented data and the best surrogate model achieved an accuracy of over 90% in predicting Defender’s classifications. The best surrogate model is then used to test four search methods to find feature values to target when modifying malicious VBA samples to evade detection. The feature values found in feature space are used to guide modification of VBA samples in code space and then tested against Defender. Over 60% of the modified malicious-samples were able to evade detection after the targeted modifications based upon the results of the best search algorithm.

Read the paper · More papers on PaperTik