Increasing Diversity with Deep Reinforcement Learning for Chatbots
Cristian Pavel, Ştefania Budulan, Traian Eugen Rebedea · 2020
Dialogue generation for open-domain conversations is a difficult and open problem that, so far, has not been able to approach human-level performance.Recently, a popular solution is to apply a sequence-to-sequence architecture, similar to the machine translation problem.These models try to map the input -given as the previous utterances, to the output -the next utterance.Unfortunately, they usually tend to repeat sentences, often preferring dull responses, that end the conversation abruptly.Therefore, Reinforcement Learning techniques have been combined with the standard sequence-to-sequence models in order to avoid their shortcomings.Our model applies a Policy Gradient method that maximizes the expected reward of generating the next utterance given a history of previous utterances.The results show an improvement in diversity up to 0.16 -almost 10x higher than the model without RL, while keeping the responses relevant to the input message.