Adversarial Attacks on Autonomous AI Agents and Mitigation Strategies

Hewa Majeed Zangana, Marwan Omar, Calvin Nobles, Anuradha Rangarajan · Advances in computational intelligence and robotics book series · 2025

Autonomous AI agents are increasingly deployed across domains such as transportation, healthcare, defense, and finance, where their decision-making capabilities enable efficiency, adaptability, and scalability. However, these systems are vulnerable to adversarial attacks that exploit weaknesses in machine learning models, perception modules, and decision policies. Such attacks can cause misclassification, system malfunction, or manipulation of agent behavior, posing significant risks to safety, reliability, and trustworthiness. This chapter explores the landscape of adversarial attacks targeting autonomous AI agents, including evasion, poisoning, and reinforcement learning–specific strategies. It further examines the implications of these threats in real-world environments where adversarial inputs may be subtle and difficult to detect. In response, a range of mitigation strategies is discussed, spanning robust model training, defensive architectures, input sanitization, anomaly detection, and hybrid approaches that combine multiple layers of defense.

Read the paper · More papers on PaperTik