Explainability versus Security: The Unintended Consequences of xAI in Cybersecurity

Marek Pawlicki, Aleksandra Pawlicka, Rafał Kozik, Michał Choraś · 2024

The rapid advancement of Artificial Intelligence in the field of cybersecurity brings about both opportunity and vulnerability, like a dual-edged sword. The research community expressed concerns over the robustness of AI against adversarial attacks, at the same time escalating the demand for transparency and accountability in the AI decision-making process. This paper highlights a critical and under-discussed paradox: the pursuit of explainability may inadvertently compromise security. The argument is that the very mechanisms which make AI decisions interpretable, such as counterexamples, can also reveal strategic insights on how to manipulate model outcomes. This paper is first to demonstrate how the Diverse Counterfactual Explanations algorithm, designed for generating counterfactual explanations, can be exploited to alter model predictions effectively. This is achieved by crafting samples tailored to flip the labels of an ML-based detector, breaching the model's integrity. The findings of this paper highlight the need for a more nuanced approach to xAI implementation in security-critical systems, one which would balance the benefits of model transparency and model robustness.

Read the paper · More papers on PaperTik