Inverse Reinforcement Learning for Multi-player Apprentice Games in Continuous-Time Nonlinear Systems
Bosen Lian, Wenqian Xue, Frank L. Lewis, Tianyou Chai, Ali Asghar Davoudi · 2021 60th IEEE Conference on Decision and Control (CDC) · 2021
We extend the inverse reinforcement learning (inverse RL) algorithms to multi-player apprentice games described by nonlinear differential equations. In these games, both the expert and the learner have N-player control inputs. Inverse RL algorithms solve the games by learner reconstructing the unknown cost function of each expert player using the demonstration of expert’s behavior (states and control inputs of each player), thereby mimicking the given behaviors. We first develop a model-based inverse RL algorithm with two learning stages: an optimal control learning stage and an inverse optimal control learning stage. Then, a model-free off-policy integral inverse RL algorithm is developed by using online expert’s demonstrations and learner’s behavior trajectories without knowing system dynamics of either expert or the learner. Finally, simulations verify the effectiveness of proposed algorithms.