Weak-to-Strong Attacking via Model Collaboration: A survey
Yameng Liu · 2025
With the widespread application of deep learning models, security in asymmetric adversarial scenarios faces severe challenges. Existing defenses primarily target "strong-to-strong" attacks, lacking effective countermeasures against Low-Resource High-Impact Attacks (LRHIA) launched through system vulnerabilities when attackers are disadvantaged in resources, model capabilities, or knowledge. This paper proposes a Weak-to-Strong Attacking via Model Collaboration (WSC) framework. Focusing on scenarios where attackers are disadvantaged in at least one dimension (resources, models, knowledge), it explores asymmetric attack paths. Through core technologies like lightweight perturbation generation, adversarial transfer of heterogeneous models, and dynamic optimization of black-box decision boundaries, the WSC framework can breach the defenses of strong models despite significant disadvantages in a single dimension. Based on a systematic analysis of cutting-edge research (2023-2025), this paper reveals the potential risks of Large Language Models (LLMs) in asymmetric confrontations and verifies that attackers can achieve efficient attacks with a disadvantage in any one dimension. The results provide a theoretical basis for designing dynamic defense mechanisms in asymmetric scenarios and lay the foundation for innovating active defense paradigms in high-security domains like edge computing.