Explainability Guided Adversarial Evasion Attacks on Malware Detectors

Kshitiz Aryal, Maanak Gupta, Mahmoud G. A. Abdelsalam, Moustafa Saleh · 2024

As the focus on security of Artificial Intelligence (AI) is becoming paramount, research on crafting and inserting optimal adversarial perturbations has become increasingly critical. In the malware domain, this adversarial sample generation relies heavily on the accuracy and placement of crafted perturbation with a goal to evade a trained classifier. This work focuses on applying explainability techniques to enhance the adversarial evasion attack on a machine-learning-based Windows PE malware detector. The explainable tool identifies the regions of PE malware files that have the most significant impact on the decision-making process of a given malware detector, and therefore, the same regions can be leveraged to inject the adversarial perturbation for maximum efficiency. Profiling all the PE malware file regions based on their impact on the malware detector's decision enables the derivation of an efficient strategy for identifying the optimal location for perturbation injection. The strategy should incorporate the region's significance in influencing the malware detector's decision and the sensitivity of the PE malware file's integrity towards modifying that region.To assess the utility of explainable AI in crafting an adversarial sample of Windows PE malware, we utilize DeepExplainer module of SHAP (SHapley Additive exPlanations) for determining the contribution of each region of PE malware to its detection by a CNN-based malware detector, MalConv. The analysis includes both local and global explanations for the given malware samples. We performed the functionality-preserving adversarial perturbation injection in different regions of PE malware wherever possible while performing non-functionality-preserving operations in a few remaining regions. This approach allows us to examine the relationship between SHAP values and the evasion rate of the adversarial attack. Furthermore, we analyzed the significance of SHAP values at a more granular level by subdividing each section of Windows PE into small subsections. We then performed an adversarial evasion attack on the subsections based on the corresponding SHAP values of the byte sequences. Our experimental evaluation shows a significant improvement in the success and efficiency of adversarial evasion attacks when injecting the perturbation in PE malware locations based on SHAP values compared to random PE locations.

Read the paper · More papers on PaperTik