Communication dans l'apprentissage robotique : utilisation du canal de tâches et social combiné pour l'apprentissage d'actions
Manuel Bied · HAL (Le Centre pour la Communication Scientifique Directe) · 2022
Creating artificial beings with human like capabilities is a long dream ofhumanity. While humanity already got closer to this goal, there is still a longway to go. A major problem preventing robots to be deployed in our day today lives is the fact that robots can not be programmed for every situationthey might encounter. This problem can be addressed by equipping robotswith the capability to learn new tasks. The ability to learn is an importanthuman social skill. Although what robots can and could accomplish focusingon manipulating their environment is already quite impressive, their socialskills can be described as, if anything, rather basic. While there are alreadyinteractive robot learning approaches (e.g. reinforcement learning and learningfrom demonstration), usually these approaches do not consider explicit teaching.Furthermore, usually they exploit actions that either accomplish the task orserve a social purpose. However, human behavior allows for actions thatcombine task and social aspects in one action. Human-Robot Interaction (HRI)has not yet sufficiently addressed these combined actions.In this thesis we focus on actions that combine task and social actions. Weare interested in how humans use these type of actions in HRI settings, howteaching influences this behavior and how these actions can be used to toaugment robot learning.We conducted a user study investigating how humans behave when teachinghow to solve a task to a robot in contrast to just solving the task. Thestudy consisted of two experiments. In the first experiments the participantswere first asked to solve a continuous maze task and to teach how to solveit to a robot afterwards. Furthermore, the participants could give negativedemonstrations that demonstrated how the task should not be solved. In thesecond experiments the demonstrations collected in the first experiment wereshown to new participants. The participants were asked how informative theyperceive these demonstrations. The results show that significantly more negativethan positive demonstrations were perceived as informative. Furthermore,significantly more demonstrations from the phase where the participants taughtto a robot were perceived as informative than from the phase where theparticipants just solved the task.We address the augmentation of robot learning with exploitation of task- andsocial channels by introducing a framework based on Reinforcement Learning(RL). In this framework we augment the reward from the environment withfeedback how an observer might perceive the actions taken by the agent.We do this by proposing three different algorithms to model the observerperception as interactive RL scheme and compare with one non-interactiveRL algorithm as baseline. In order to model the observer we vary the methodhow the observer estimates how likely the agent is going for the real goal.We evaluate our approach on five environments and calculate the legibilityof the learned trajectories. Legibility is a scalar metric measuring how wellgoals can be inferred from actions. The results show that the legibility of thelearned trajectories is significantly higher while integrating the feedback fromthe observer compared with a standard Q-Learning algorithm not using theobserver feedback.From these results we conclude that humans use combined actions in HRIsettings to enrich the communication, but also perceive these actions as infor-mative. Further, that combined actions can be learned with a RL frameworkby integrating reasoning about potential observers to enrich the actions withsocial aspects. While the research presented in thesis is limited to specificcases, it demonstrates the promising potential of combined actions in HRIsettings.