Group Confident Policy Optimization

Yao Li, Zhenglin Liang · 2025

Learning-based policy improvement methods rely on extensive data collection and iterative training, leading to credibility challenges during early training stage when the agent is transferred to new, uncertain and risky environments. To address this, this paper develops a novel reinforcement learning algorithm, termed Group Confident Policy Optimization (GCPO), which emphasizes enhancing the safety and confidence of exploration processes and policy updates. The algorithm proposes a confident advantage function that leverages group normalization to mitigate training variance induced by sampled data and introduces a confidence-based corrective factor. By employing confidence-augmented policy gradient updates, this method ensures safe agent behaviors and progressive performance improvement throughout the training cycle. Simulation experiments demonstrate that GCPO achieves superior performance in risky tasks compared to the traditional PPO method, while the ablation study verifies the importance of confident advantage estimation. These findings contribute to establishing trustworthy engineering paradigms for safety-critical automation in cross-environment transfer scenarios.

Read the paper · More papers on PaperTik