Dynamic risk mitigation in edge computing: A queuing-theoretic framework for reliability-important cascading failure prevention
Kadir Sarikaya · Reliability Engineering & System Safety · 2026
Edge computing infrastructure is increasingly deployed in reliability-important domains but remains vulnerable to cascading failures driven by load redistribution. This paper presents a queue-motivated, graph-aware continuous-control framework that combines a proxy-informed load-dependent hazard model with a Graph Attention Soft Actor–Critic (GAT-SAC) policy for graceful throttling. The hazard model is anchored by a proxy-based calibration on 334 retained rows from a 358-incident source dataset and is used as an uncertainty-aware working slope estimate rather than as a precise physical identification result. The controller is evaluated under unified load-shedding physics in a fixed high-stress BA-100 regime using a frozen bank of 40 paired scenarios. In this canonical benchmark, GAT-SAC achieves the highest average system survival time (14.03 h), outperforming No Defence (4.50 h), redundancy baselines, and the threshold-based heuristic baseline. Sensitivity analysis identifies a cliff effect near ϕ ≈ 0 . 05 , where modest load shedding shifts the system from a cascade-dominated regime to a stabilized regime. End-to-end inference latency is 0.81 ms on laptop-class CPU hardware, which is negligible relative to the 1-minute control interval. Broader topology and size-transfer checks are reported only as supplementary descriptive probes.