Reinforcement learning of LQR control policy by a double inverted-pendulum biomechanical model

Kamran Iqbal, Muhammad Haras · 2023

Optimal LQR feedback gains can be learned using reinforcement learning (RL) framework for systems with unknown dynamics using policy iteration methods. However, policy iteration in the case of inherently unstable systems becomes challenging. In this study we establish reinforcement learning of optimal feedback gains in the case of a nonlinear double inverted-pendulum (DIP) biomechanical model. Using an admissible initial policy, the biomechanical model was simulated in MATLAB and trajectory data were recorded. The state variables were transformed to quadratic basis function and used in approximate dynamic programming (ADP) to learn the solution to the algebraic Riccati equation (ARE) underlying the LQR problem. The RL results obtained in the case of an inherently unstable DIP system indicate relatively fast convergence and demonstrate the potential to apply RL techniques to more complex systems.

Read the paper · More papers on PaperTik