Online Monitoring of High-Dimensional Data Streams With Deep Q-Network
Haoqian Li, Ziqian Zheng, Kaibo Liu · IEEE Transactions on Automation Science and Engineering · 2025
With the fast advancements in Internet of Things (IoT) technology and sensing infrastructure, a wide range of systems continues to generate a massive amount of data. Meanwhile, practical resource constraints such as limited bandwidth or processing capability restrict the full observability of data streams in real time. As a result, the practitioners often need to dynamically decide which data streams to observe given the resource constraints in order to quickly detect any system anomaly as soon as possible. In this article, we propose a reinforcement learning framework based on deep Q-learning and combine it with statistical process control (SPC) techniques to effectively monitor high-dimensional data streams when only partial observations are available at each acquisition time due to resource constraints. To the best of our knowledge, this is the first work that integrates deep reinforcement learning, which considers long-term rewards associated with the dynamic data sampling strategy, into SPC charts; this integration effectively addresses the challenge of resource constraints in monitoring high-dimensional data streams. Specifically, we construct a nonparametric monitoring statistic for each data stream and develop a reinforcement learning framework to automatically identify the most informative data streams for observation at each time epoch. The state space, action space, and rewards in the reinforcement learning framework are carefully designed and a Double Dueling Q-network is trained accordingly. Unlike existing methods, which rely on heuristic approaches to determine the sampling strategy, the proposed framework maximizes the long-term reward, thus leading to superior performance compared to existing benchmarks. Numerical simulations and a case study are thoroughly conducted, showing that the proposed method outperforms the state-of-the-art algorithms by significantly reducing detection delay. Note to Practitioners—This paper is motivated by the practical issue of online process monitoring and anomaly detection with resource constraints. In particular, due to resource constraints, practitioners can only select a subset of data streams to monitor at each time epoch. Thus, the central challenges are to dynamically choose which data streams to observe and to decide when to raise an alarm. Unlike the existing methodologies which are heuristic and only consider short-term rewards from the dynamic sampling, this paper proposes a novel reinforcement learning framework that allows practitioners to monitor high-dimensional heterogeneous data streams more efficiently by considering the long-term rewards associated with the dynamic data sampling strategy. Four main steps are involved in the proposed method: (i) construct a nonparametric local monitoring statistic for each data stream; (ii) train the Double Dueling Q-network offline according to the proposed deep reinforcement learning framework; (iii) at each time epoch of the online monitoring, determine the most informative data streams to observe based on the deep Q-network; and (iv) determine whether to raise an alarm in the system.