Surrogate-Guided Adversarial Attacks: Enabling White-Box Methods in Black-Box Scenarios
Dimitrios Christos Asimopoulos, Panagiotis I. Radoglou Grammatikis, Panagiotis Fouliras, Konstandinos Panitsidis, Georgios Efstathopoulos, Θωμάς Λάγκας, Vasileios Argyriou, Igor Kotsiuba, Panagiotis G. Sarigiannidis · 2025
Adversarial attacks pose significant threats to machine learning models, with white-box attacks such as Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), and Basic Iterative Method (BIM) achieving high success rates when model gradients are accessible. However, in real-world scenarios, direct access to model internals is often restricted, necessitating black-box attack strategies that typically suffer from lower effectiveness. In this work, we propose a novel approach to transform white-box attacks into black-box attacks by leveraging state-of-the-art surrogate models, including MultiLayer Perceptrons (MLP) and XGBoost (XGB). Our method involves training a surrogate model to mimic the decision boundaries of an inaccessible target model using pseudo-labeling, thereby enabling the application of gradient-based white-box attacks in a black-box setting. We systematically compare our approach against conventional black-box attacks, such as Zero Order Optimization (ZOO), evaluating their effectiveness in terms of attack success rates, transferability, and computational efficiency. The results demonstrate that surrogate-assisted attacks perform as good as standard black-box methods, bridging the performance gap between white-box and black-box adversarial attacks. This study highlights the power of surrogate models in enhancing adversarial transferability and provides insights into the robustness of different machine learning architectures against adversarial threats.