Adaptive Intrusion Detection with PPO and Zero Trust Continuous Verification

Ekramul Haque, Kamrul Hasan, Imtiaz Ahmed, Sharif Ullah, Md. Tariqul Islam, Mir Mehedi Ahsan Pritom · 2025

Cyber threats have evolved rapidly, and traditional rule-based intrusion detection systems mostly fail to adapt to emerging attack patterns. Our approach combines Reinforcement Learning (RL) with Proximal Policy Optimization (PPO) in an intrusion detection system (IDS) framework implementing continuous verification, a core principle of Zero Trust Architecture (ZTA). Our approach allows for real-time monitoring through behavioral analysis: a fix-or-detect-then-defend strategy against a given threat without making assumptions of prior trust. We trained the PPO model on the CIC-IDS2017 and CIC-IDS2018 datasets and achieved a max 99.19% accuracy and 97.97% F1 score in traffic classification. For greater transparency, we applied Shapley Additive Explanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) to provide clear insights into the model's decisions. The powerful fusion of RL, ZTA, and explainable AI (XAI) results in an adaptive, interpretable, and resilient IDS that can counter sophisticated cyber threats. In some sense, our work is the first to merge RL-driven adaptive security with continuous verification, thus advancing the robustness and transparency of modern IDS solutions.

Read the paper · More papers on PaperTik