Model-based Safe Reinforcement Learning using Variable Horizon Rollouts
Subrata Roy Gupta, Utkarsh Suryaman, Rahul Narava, Shashi Shekhar Jha · 2024
Safe reinforcement learning aims to ensure the safety of agents and their interactions with the environment. Traditional reinforcement learning algorithms often neglect safety considerations, resulting in undesirable consequences when deployed in real-world scenarios. To address this issue, safe reinforcement learning algorithms incorporate safety constraints or modify the reward distribution to prioritize the safety of the agent. Recent literature on model-based safe reinforcement learning uses a fixed look-ahead horizon to avoid unsafe states, limiting the agent’s adaptability to changing environments. In this paper, we propose a variable horizon look-ahead based on the agent’s current state, resulting in improved performance and adaptability in uncertain environments. Our approach leverages the Gaussian process regression model to estimate the value of the horizon dynamically based on the agent’s current state. To enhance sample efficiency, we introduce a selective sampling strategy that reduces the number of rollouts by eliminating samples that may lead to unsafe states and also optimizes the use of computational resources while ensuring safety. We evaluate the effectiveness of our approach in multiple mujoco environments. Our results demonstrate a significant reduction in terms of safety violations during validating compared to existing approaches from the literature.