Surrogate-Guided Adversarial Attacks: Enabling White-Box Methods in Black-Box Scenarios

Dimitrios Christos Asimopoulos, Panagiotis I. Radoglou Grammatikis, Panagiotis Fouliras, Konstandinos Panitsidis, Georgios Efstathopoulos, Θωμάς Λάγκας, Vasileios Argyriou, Igor Kotsiuba, Panagiotis G. Sarigiannidis · 2025

Adversarial attacks pose significant threats to machine learning models, with white-box attacks such as Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), and Basic Iterative Method (BIM) achieving high success rates when model gradients are accessible. However, in real-world scenarios, direct access to model internals is often restricted, necessitating black-box attack strategies that typically suffer from lower effectiveness. In this work, we propose a novel approach to transform white-box attacks into black-box attacks by leveraging state-of-the-art surrogate models, including MultiLayer Perceptrons (MLP) and XGBoost (XGB). Our method involves training a surrogate model to mimic the decision boundaries of an inaccessible target model using pseudo-labeling, thereby enabling the application of gradient-based white-box attacks in a black-box setting. We systematically compare our approach against conventional black-box attacks, such as Zero Order Optimization (ZOO), evaluating their effectiveness in terms of attack success rates, transferability, and computational efficiency. The results demonstrate that surrogate-assisted attacks perform as good as standard black-box methods, bridging the performance gap between white-box and black-box adversarial attacks. This study highlights the power of surrogate models in enhancing adversarial transferability and provides insights into the robustness of different machine learning architectures against adversarial threats.

Read the paper · More papers on PaperTik