Evaluating Reinforcement Learning Reward Functions for APT Detection in Industrial IoT Systems

Mohammad Alauthman, Amjad Yousef Aldweesh, Ahmad Al–Qerem, Ibtisal Daoud, Mouhammd Sharari Alkasassbeh, Amjad Gawanmeh · 2025

Advanced Persistent Threats (APTs) pose significant challenges in Industrial Internet of Things (IIoT) environments, as they often employ stealthy and multi-stage attack strategies. The complexity of heterogeneous data sources, including provenance and network logs, further complicates detection. In this paper, we leverage the newly introduced CICAPT-IIoT2024 dataset as a benchmark for evaluating Intrusion Detection Systems (IDS). We focus on three Reinforcement Learning (RL) reward functions—Weighted Precision and Recall, Cost-Sensitive, and Balanced Accuracy—to determine their effectiveness in detecting APTs. Our experiments show that selecting the right reward function greatly influences detection trade-offs such as precision, recall, and false positive rates. The findings highlight the nuanced role of reward tuning and emphasize that a cost-sensitive approach more effectively penalizes false negatives in high-stakes IIoT scenarios, whereas balanced accuracy rewards maintain a more uniform trade-off between detection and false alarms. We discuss the implications of these outcomes for real-world IIoT deployments and propose directions for future research in improving RL-based IDS performance.

Read the paper · More papers on PaperTik