Exploring Adversarial Attacks in Software Analytics through Machine Learning Explainability
Anonymous · Zenodo (CERN European Organization for Nuclear Research) · 2023
In recent years, machine learning (ML) models have been extensively used in software analytics, such as code completion, malware detection, code clone detection, code authorship attribution, code search, API recommendation, and code comment generation. However, studies show that state-of-the-art ML models are vulnerable to adversarial attacks when we add minimal perturbations to the original input, which may lead to technical debt and substantial monetary losses in software analytics. As a result, the ML models’ robustness against adversarial attacks must be assessed before they are deployed in software analytics. ML explainability has recently gained popularity and gives us insights into the reasoning behind the ML models’ predictions. This study aims to investigate the relationship between ML explainability and adversarial attacks. Furthermore, we are interested in generating adversarial examples based on the explanation provided by the ML explainability techniques to measure the robustness of the ML models. In this paper, we select five datasets, three ML explainability techniques, and seven state-of-the-art ML models to conduct our study. Our large-scale experimental results demonstrate a positive correlation between ML explainability techniques and adversarial attacks. Furthermore, modifying the feature values of the top-𝑘 important features identified by ML explainability can generate effective adversarial examples. Therefore, the ML models under attack fail to accurately predict up to 86.6% of instances that were correctly predicted before adversarial attacks, indicating the models’ low robustness against such attacks