The Intelligent Trap: Markov Decision Process Framework for Dynamic Honeypot
Maryam Var Naseri · 2025
Modern adversarial attacks are integrated with AI to target and evade network security solutions such as honeypots. Honeypots aim to learn from attackers’ behaviour. However, most current honeypots are configured and managed statically, wherein prior knowledge about the attackers is mandatory. The current intelligent honeypots, such as HARM proposed by Dowling, have shown limitations in adapting to the evolving automated and repetitive malware attacks, especially those targeting the Internet of Things. Furthermore, QRASSH is another example of an autonomic honeypot designed to address the issue of inadaptability to evolving threats. However, the developers of this project encountered difficulties in integrating machine learning techniques and in formulating effective reward functions for reinforcement learning. Moreover, Heliza faced multiple challenges. These included difficulties in identifying attack types and distinguishing between human-based and automated attacks. The system also struggled to define optimal learning strategies and limitations to adapt to the techniques employed by attackers. In response to these challenges, the main objective of this research is to design a honeypot with autonomous behaviour, enabling the system to make decisions in a dynamic network environment. The first objective of this research is to define the concept of an autonomic honeypot by exploring current autonomic honeypot systems. The second objective is to design a Markov Decision-Making Process (MDP) model to better understand the attackers' actions in different situations. The third objective involves deploying the MDP model for honeypot integration using reinforcement learning algorithms, allowing the honeypot to generate responses while interacting with new attackers autonomously. Finally, the fourth objective focuses on the evaluation of the honeypot's performance in terms of adaptability and learning stability. The first objective of this research focuses on a new class of intelligent honeypots; however, their adaptability characteristics offer greater protection against attackers' detection techniques. This research reviewed the current intelligent honeypot systems using criteria based on the Autonomic Computing Toolkit User’s Guide definition of autonomic systems. This research describes an autonomic honeypot as a system that anticipates attacker behaviour, protects data, supports dynamic response generation and autonomous decision-making capabilities. In the second objective, various experimental studies were conducted in a controlled environment to evaluate the honeypot's behaviour before integration with a modified configuration, and after integration. The first round of experiments undertaken was to focus on identifying attackers’ goals and techniques, and capturing the normal behaviour of attackers. This data allowed us to create a Markov Decision-Making Process (MDP) model for the integration round of the experiments and evaluate Cowrie's adaptive behaviour while interacting with attackers. The MDP model helped to measure the probability of attackers taking action in different situations. This enhances the honeypot with active intelligence-gathering for attackers' unique pattern recognition. The third objective investigated whether an MDP-based model could offer adaptability, intelligent gathering, and dynamic response features while interacting with the attackers. By deploying reinforcement learning techniques, the model achieved significant improvement in identifying unique patterns, adaptability, and dynamic response generation for cryptomining and botnet attack scenarios. The fourth and last objective evaluated the MDP-based honeypot regarding its adaptability in response generation and learning stability toward new evolving threats after its integration.