Reinforcement Learning-Guided Adversarial Training to Enhance Intrusion Detection Robustness on the Mirai Dataset

Hangrui Hu · Applied and Computational Engineering · 2025

Despite demonstrating impressive detection capabilities, deep learning-based intrusion detection systems (IDS) exhibit significant vulnerability to adversarial examples, particularly under gradient-based attacks such as FGSM. To address this issue, this study proposes a reinforcement learning (RL)-guided adversarial training framework, in which a Q-learning agent adaptively optimizes perturbation strengths (ε values) during training. This adaptive mechanism allows the model to defend against a broad range of threat intensities without manual tuning, thereby overcoming the limitations of traditional fixed-ε adversarial training methods. The proposed approach is evaluated on the Mirai subset of the Kitsune dataset, selecting 40,000 labeled samples. Experimental results demonstrate that the RL-based adversarial training model significantly outperforms both a baseline model (trained only on clean data) and a fixed-ε model (ε = 0.1) in terms of robustness. For instance, under strong adversarial perturbation (ε = 0.2), the RL-guided model achieves an accuracy of 96%, compared to 90% for the fixed-ε model and 82% for the baseline. Furthermore, the agent converges toward selecting ε ≈ 0.11 as the optimal trade-off between training stability and defense effectiveness. These findings validate the efficacy of reinforcement learning in dynamically optimizing adversarial training, offering a scalable defense against gradient-based attacks.

Read the paper · More papers on PaperTik