Reinforcement learning of dialogue strategies using the user's last dialogue act
Matthew Frampton, Oliver Lemon · 2005
Previous attempts at using reinforcement learning to design dialogue strategies for spoken dialogue systems e.g. [Singh et al., 2002; Pietquin and Re-nals, 2002] have included only ‘low-level ’ informa-tion in state representations i.e. whether or not a slot (e.g. destination city) has been filled and the con-fidence score associated with any supplied value. We explore the benefits of adding limited ‘high-level ’ contextual information, in this case the di-alogue act of the last user utterance. A general concern with adding more information is that the size of the state space might increase to a degree where learning becomes intractable. We describe 3 experiments in this paper which involve learn-ing dialogue strategies for a flight booking dia-logue system, and which test the potential benefits of adding this ‘high-level ’ information to the state representation. We also explore the use of differ-ent reward functions in learning. In the first ex-periment, adding the high level information does not result in a superior learned strategy. However, the second experiment demonstrates a first simple case in which including the ‘high-level ’ informa-tion does result in a superior learned strategy, pro-ducing a 52 % increase in average reward, and the third scales up the problem to 4 slots. Here a reward function that rewards only totally correct database queries (rather than also valuing partial correctness) is found to produce the best strategy. 1