Adversarial Reinforcement Learning Based IoT Honeypot

Hao Zhang, Siyuan Zhang, Chengrun He, Chengcheng Zhao · 2025

Internet of Things (IoT) honeypots are decoy systems deployed to entice attackers to gather threat intelligence and protect real systems. High-interaction IoT honeypots powered by reinforcement learning (RL) have emerged as a promising solution due to their cost-effectiveness and scalability. However, these systems are typically based on the assumption that attackers exhibit stationary behavior. In reality, attack strategies against IoT can be dynamic and adaptive, creating a non-stationary environment due to the adversarial nature of attacker-honeypot interactions. To solve this issue, we propose an IoT honeypot based on adversarial reinforcement learning, i.e., Repeated-Update-Q-learning (RUQL, a classical RL method for non-stationary environments). It is composed of a data preprocessing module, an RUQL module, and a response database. Experimental results show that compared to honeypots based on random strategies, classical Q-learning, and deep RL, the proposed system can effectively respond to attacks and improve attack capture and analysis capabilities.

Read the paper · More papers on PaperTik