Pessimistic policy iteration with bounded uncertainty
Zhiyong Peng, Changlin Han, Yadong Liu, Jingsheng Tang, Zongtan Zhou · Expert Systems with Applications · 2025
Offline Reinforcement Learning (RL) aims to learn policies by using static datasets. The extrapolation error in out-of-distribution (OOD) samples can cause off-policy RL algorithms to perform poorly on offline datasets. Hence, it is critical to avoid visiting OOD states and taking OOD actions in offline RL. Several recent methods have used uncertainty estimation to distinguish OOD samples. However, errors in the uncertainty estimation make the purely uncertainty-based method unstable and require additional components to ensure sufficient pessimism . In this study, we propose a Bounded Uncertainty based Pessimistic policy iteration algorithm (BUP). The BUP pessimistically estimates the value function via bounded uncertainty, and the uncertainty bound is achieved by constraining the actor from taking highly uncertain actions. The suboptimality bound of BUP is theoretically guaranteed in linear Markov Decision Processes (MDPs), and experiments on D4RL datasets show that BUP matches the state-of-the-art performance. Moreover, BUP is simple to implement with low computational cost and does not require any additional components.