Multi-agent deep reinforcement learning for penetration testing of IoT devices through their mobile companion app
Francesco Pagano, Mariano Ceccato, Alessio Merlo, Paolo Tonella · Journal of Systems and Software · 2026
The increasing integration of IoT devices into critical infrastructure has made them prime targets for cyberattacks. Many of these devices rely on outdated or legacy software, which introduces inherent vulnerabilities and complicates firmware updates, making identifying and testing these weaknesses essential. Traditional methods typically employ black-box approaches, mutating network requests generated during device operation to craft potential attack vectors. However, these methods face limitations when dealing with encrypted or proprietary protocols. Recent tools, such as Diane and IoTFuzzer, interact with IoT devices through their mobile companion apps and use fuzzing techniques to modify request content, causing crashes in IoT device software. Although these approaches can effectively trigger software crashes, they do not generate actual exploits, as they do not precisely target or exploit specific vulnerabilities. To address these limitations, we introduce MITHRAS, the first approach that uses mobile companion apps to deliver maliciously mutated requests directly to IoT devices, explicitly targeting Remote Code Execution (RCE) vulnerabilities. MITHRAS uses Deep Reinforcement Learning to efficiently navigate the communication code within companion apps, dynamically mutating request payloads before transmission. Adapting to previous attack outcomes, MITHRAS refines its strategy, mimicking human decision-making to improve the effectiveness of exploit generation.