Online Training and Pruning of Deep Reinforcement Learning Networks
Valentin Frank Ingmar Guenter, Athanasios Sideris · IET Cyber-Systems and Robotics · 2026
ABSTRACT Scaling deep neural networks (NNs) of reinforcement learning (RL) algorithms has been shown to enhance performance when feature extraction networks are used, but the gained performance comes at the significant expense of increased computational and memory complexity. NN pruning methods have successfully addressed this challenge in supervised learning, but have only recently been considered in RL applications. We propose an approach to integrate simultaneous training and pruning within advanced RL methods, in particular, RL algorithms enhanced by the online feature extractor network (OFENet). Our networks (XiNet) are trained to solve stochastic optimisation problems over the weights of the RL networks and the parameters of variational Bernoulli distributions for 0/1 random variables (RVs) scaling the th unit in these networks. The stochastic problem formulation induces regularisation terms that promote convergence of the variational parameters to 0 when a unit contributes little to the performance. In this case, the corresponding structure is rendered permanently inactive and pruned from its network. We express the parameter and computational complexity of the RL networks, and in particular, the DenseNet networks involved in the architecture of OFENets, in terms of the parameters of their RVs and obtain a complexity‐aware cost. We then match this cost with the pruning‐promoting regularisation terms in the stochastic optimisation problem and show that many hyperparameters associated with them can be automatically selected in terms of one user‐defined hyperparameter for each network. Thus, the RL objectives and network compression are effectively combined. We evaluate our method on continuous control benchmarks (MuJoCo) and popular RL agents, demonstrating that OFENets can be pruned considerably with minimal loss in performance. Furthermore, our results suggest that pruning larger networks during training produces more efficient and higher‐performing RL agents than training smaller networks from scratch.