Learning Landmark-Oriented Subgoals for Visual Navigation Using Trajectory Memory
Jia Qu, Shotaro Miwa, Yukiyasu Domae · 2022 IEEE Symposium Series on Computational Intelligence (SSCI) · 2022
In many deep reinforcement learning (DRL) applications, agents must perform complex and long-horizon tasks that are still challenging in DRL because of temporally ex-tended tasks with sparse rewards. Goal-conditioned hierarchical reinforcement learning (HRL) is a promising approach for control at multiple time scales via a hierarchical structure with subgoals. One of the key issues of goal-conditioned HRL is the definition of subgoals. In this study, we propose a DRL model for learning landmark-oriented subgoals using attention-augmented trajectory memory. In our approach the agent is trained to make decisions based on both 1) current perception, which is a short-term temporal history derived from vanilla long short-term memory (LSTM) and 2) trajectory memory, which represents a contextual summary of long-term historical LSTM states augmented by attention. The experiment on a visual navigation task shows that the short-term LSTM state of the current perception module extracts landmark subgoals as clusters that correspond to lower-level policies, and the long-term context states of trajectory memory extract subgoal transitions that correspond to higher-level policies. Furthermore, the proposed method demonstrated superior adaptability to environmental changes,