“How Does It Detect A Malicious App?” Explaining the Predictions of AI-based Malware Detector
Zhi Lü, Vrizlynn L. L. Thing · 2022
In this paper, we present a novel model-agnostic explanation method for AI-based malware detection models, to assure the trustworthiness from both cyber security and AI practitioners' perspectives. Our proposed method identifies and quantifies the data features relevant to the predictions by two steps: i) data perturbation that generates the synthetic data by manipulating features' values; and ii) optimization of features attribution values to seek significant changes of prediction scores on the perturbed data with minimal feature values changes. The proposed method is validated by three experiments. We firstly demonstrate that our proposed model explanation method can aid in discovering how AI models are evaded by adversarial samples quantitatively. In the following experiments, we compare the explainability and fidelity of our proposed method with state-of-the-arts. The results show that the proposed method can explain different AI models with robust explanations and high fidelity.