Subgoal Discovery for Hierarchical Dialogue Policy Learning
Da Quan Tang, Xiujun Li, Jianfeng Gao, Chong Wang, Lihong Li, Tony Jebara · 2018
Developing agents to engage in complex goaloriented dialogues is challenging partly because the main learning signals are very sparse in long conversations.In this paper, we propose a divide-and-conquer approach that discovers and exploits the hidden structure of the task to enable efficient policy learning.First, given successful example dialogues, we propose the Subgoal Discovery Network (SDN) to divide a complex goal-oriented task into a set of simpler subgoals in an unsupervised fashion.We then use these subgoals to learn a multi-level policy by hierarchical reinforcement learning.We demonstrate our method by building a dialogue agent for the composite task of travel planning.Experiments with simulated and real users show that our approach performs competitively against a state-of-theart method that requires human-defined subgoals.Moreover, we show that the learned subgoals are often human comprehensible.