Reward Function Design for Stand-Off Tracking of Reinforcement Learning

Yeontaek Jung, Jinrae Kim, Seong-hun Kim, Youdan Kim · AIAA SCITECH 2023 Forum · 2023

View Video Presentation: https://doi.org/10.2514/6.2023-1440.vid As a reward function encodes the objective of reinforcement learning, the overall control performance and convergence rate of the learning can be adjusted by reward function design. In this study, new reward functions are proposed for control using reinforcement learning. With an attempt to utilize Lyapunov stability theory, the proposed reward functions are needed to set Lyapunov candidate function. The reward is computed depending on whether the variation of the Lyapunov candidate function is negative or positive. As a result, the reinforcement learning agent learns which region of state is good or bad. With exponential function whose exponent is a quadratic function, the agent is also guided to learn how to get the good region. The stability of the proposed reward functions is analyzed. To verify the performance of the reward function, numerical simulation is performed for stand-off tracking using Proximal Policy Optimization algorithm. The results show that the proposed reward functions have improved regulation performance and data efficiency.

Read the paper · More papers on PaperTik