DRLFailureMonitor: A Dynamic Failure Monitoring Approach for Deep Reinforcement Learning System

Yi Cai, Xiaohui Wan, Zhihao Liu, Zheng Zheng · 2024

Over the past decade, deep reinforcement learning (DRL) has seen increasing adoption in addressing various sequential decision-making tasks, such as autonomous driving and robotic control, demonstrating superior performance. However, as its application continues to broaden, the reliability of these DRL systems encounters significant challenges, particularly within safety-critical domains where any system failure could lead to catastrophic consequences. Currently, the assurance of reliability in DRL systems relies on testing techniques, which consist of offline solutions that uncover and address potential defects before deployment. In contrast to existing studies, this paper addresses the issue of online monitoring of DRL systems to detect and alert potential failures in advance, thereby facilitating the transition from automated to manual decision-making when necessary. Specifically, we model online monitoring of DRL systems as a multivariate time series classification problem and propose a novel failure monitoring approach, which is named DRLFailureMonitor. This method employs a temporal dynamic graph neural network to capture hidden spatiotemporal dependencies in the sequences of trajectories planned by the DRL systems. Extensive experiments across five benchmark DRL environments demonstrate that DRLFailureMonitor achieves an average failure detection accuracy of 98.2% and a recall of 100%. In addition, the proposed method offers a failure detection lead time ranging from 11 to 79 steps, indicating that it can detect failures in DRL systems well before the system actually fails and causes any losses. Consequently, this method holds significant importance for enhancing human-machine collaboration and improving the reliability of DRL systems in safety-critical areas.

Read the paper · More papers on PaperTik