GCB-PPO2: A Hybrid Deep Reinforcement Learning Intrusion Detection System for Under-Represented Attack Categories in SDN
Jue Chen, Tao Hongyu, Cui Meng, Peng Haidong, Xihe Qiu · IEEE Transactions on Network Science and Engineering · 2025
The centralized control plane inherent in SDN architecture creates critical security dependencies, where malicious exploitation of controller vulnerabilities could propagate systemic network failures. The Intrusion Detection System (IDS) effectively counters cyber threats, and Deep Reinforcement Learning (DRL) can enable the IDS to dynamically adapt to the constantly evolving attack patterns through autonomous environmental interaction and real-time policy optimization. However, the current DRL based IDS face the limitations of minority-class detection, generalization gap, and tuning complexity. This paper propose GCB-PPO2 (Generative Adversarial Networks CNNA-BiLSTM Proximal Policy Optimization 2), a hybrid DRL system synergistically combining Generative Adversarial Networks (GAN) and Proximal Policy Optimization 2 (PPO2). This framework integrates three innovations: (1) Unlike prior DRL models cannot detect low-frequency attacks correctly, we design a C-GAN (Conditional Generative Adversarial Networks) architecture to generate targeted under-represented attack samples for the SDN scenario intentionally, improving the recognition accuracy of the model for minority classes significantly. (2) Unlike existing DRL models that rely on single-modal networks and experiment on single datasets, we propose a PPO2-based DRL framework with a CNN-LSTM shared network to optimize dynamic policy adaptation while capturing spatio-temporal patterns, enhancing generalization ability to different environmental dynamics. (3) Unlike conventional DRL implementations that manually tune parameters or use random search, we embed a Bayesian optimization as a dedicated hyper-parameter auto-tuner, systematically resolving sensitivity bottlenecks and ensuring robustness in fluctuating environments. Experimental results demonstrate that GCB-PPO2 model achieves remarkable accuracy of 99.92% and 99.01% for binary and multiple classification scenarios, respectively, with F1-scores exceeding 85.91% for under-represented attack categories in InSDN dataset. Moreover, the model proves its strong generalization capability for maintaining over 98.65% accuracy on another dataset, while confirming its improved training efficiency for surpassing conventional hybrid deep learning models by more than 22.27% in training time.