Decomposing Control Lyapunov Functions for Efficient Reinforcement Learning
Antonio Manuel López, David Fridovich-Keil · 2025
Recent methods using Reinforcement Learning (RL) have proven to be successful for training intelligent agents in unknown environments. However, current state-of-the-art RL methods require large amounts of data to learn a specific task, leading to unreasonable costs when deploying the agent to collect data in real-world applications. In this paper, we take a control-theoretic approach to improve sample efficiency in RL. Our approach has two key components. First, we build upon recent work which shows that the inclusion of a Control Lyapunov Function (CLF) in the reward function can enable RL algorithms to learn with substantially lower discount factors, and therefore require less data to train. To do this “reward shaping” requires access to a CLF, however, and computational methods to compute a CLF do not scale well to high-dimensional systems. Therefore, the primary contribution of our work is the development of a new variety of CLF-like functions, which we term Decomposed Control Lyapunov Functions (DCLFs). We show that these functions can be used in the same way as CLFs for RL reward shaping, yet are more readily computable in higher-dimensional cases via a system decomposition technique. Through multiple examples, we demonstrate the effectiveness of this approach, where our method finds a policy to successfully land a quadcopter with less than half the amount of real-world data required by two state-of-the-art RL algorithms.