Off-Policy Reinforcement Learning for Optimal Control of a Two Wheeled Self Balancing Robot
Athira Mullachery, Shaikshavali Chitraganti · 2023
We consider a model free off-policy based rein-forcement learning algorithm developed for the linear quadratic regulator control of a two wheeled self balancing robot (TWSBR) in discrete time setting. The usual methods to control TWSBR utilises its mathematical model that may not always be available or may not be accurate. Reinforcement learning (RL), which is an important branch of artificial intelligence, offers model free techniques that can be applied to systems whose complete model is unknown. The proposed approach uses a subtopic of reinforcement learning called as off-policy RL, employing separate policies for data generation and value function evaluation. The proposed model free off-policy RL algorithm improves the control policy using generated data samples and offers the advantage of being immune to bias arising from probing noise added to the control input to satisfy the requirement of persistence of excitation. Numerical simulations are provided to demonstrate effectiveness of the approach.