Reinforcement Learning to Efficiently Recover Control Performance of Robots Using Imitation Learning After Failure
Shoki Kobayashi, Takeshi Shibuya · 2022 IEEE International Conference on Systems, Man, and Cybernetics (SMC) · 2022
Extreme environments, such as space and underwater, are difficult for humans to enter because they involve risks, hence it is necessary to employ autonomous robots instead of humans. When robots fail in extreme environments, it is essential for the robot to automatically recover control following the rules of failure because humans cannot repair the robot directly. Reinforcement learning is expected to automatically acquire the control rule; however, the retrieval of the control rule requires significant trial-and-error. Imitation learning cannot acquire the control rule if no suitable expert data exist. Methods combining imitation learning and reinforcement learning reduce the number of trial-and-errors; however, they are still not effective against robot failure because these methods cannot utilize expert data effectively. This paper proposes a reinforcement learning method that efficiently recovers control performance from failures by utilizing both the control rules prepared by the designer and multiple discriminators to calculate measures similar to expert data. Experimental results show that the proposed method recovers the control performance with fewer episodes than the conventional method. The main contribution of the proposed method is its efficiency against robot failure through utilizing expart data prepared for failure by the designers for imitation learning.