CTR-Net: Scalable safe reinforcement learning via neural approximations of control theoretic regulators

Ramen Ghosh · Intelligent Systems with Applications · 2025

Ensuring hard constraint satisfaction during both training and deployment is central to safety–critical reinforcement learning (RL). Control–theoretic regularization (CTR) enforces safety by filtering actions through viability- or barrier–certified safe sets, but evaluating the state–dependent regulator R ( x ) online is often prohibitive in high dimensions. We propose a scalable CTR framework based on neural regulator approximators R ˆ θ ( x ) —differentiable surrogates of R ( x ) that enable fast projection or rejection–sampling filters within standard RL loops. We formalize a learning–theoretic analysis for approximate safety filtering and prove probably approximately correct (PAC)–style guarantees: if the set approximation error is bounded by ɛ with confidence 1 − δ , then the probability of constraint violation along a length– T rollout is bounded by a term that scales linearly in T and ɛ (plus δ ). We further show that the performance suboptimality of the filtered policy is controlled analytically by the same approximation envelope, yielding an explicit, provably quantified safety–versus–optimality tradeoff (PAC bounds linear in T and the envelope), complemented by empirical ablations; see also the calculus-of-variations view of constrained tradeoffs (Younis, 2023). The resulting method, CTR-Net , is architecture–agnostic and supports real–time execution via fast, differentiable safety layers. Empirical evaluations on high–dimensional continuous–control benchmarks—including safe locomotion and constrained multi–joint manipulation—demonstrate reliable constraint satisfaction during learning and deployment, robustness under modeling uncertainty and substantial computational gains relative to exact viability/barrier baselines. By coupling operator–free neural safety sets with CTR guarantees, CTR-Net bridges theoretical safety certificates and scalable implementation, advancing practical, real–time safe RL for complex intelligent systems. • CTR-Net learns neural approximations of control-theoretic safety regulators. • Enforces hard safety during training via fast projection or rejection. • Provides PAC safety and performance bounds under approximation error. • Achieves near-zero violations and competitive return on MuJoCo tasks. • Real-time feasible: sub-millisecond safety filtering at 50–100 Hz.

Read the paper · More papers on PaperTik