Utterance Censorship of Online Reinforcement Learning Chatbot
Yixuan Chai, Guohua Liu · 2018
Researchers have applied online deep reinforcement learning in order to enhance the open-domain conversational skills of chatbots. These chatbots have the ability to learn conversations from real users but in practical applications, some users may take advantage of the chatbot's online learning ability to generate offensive responses. In this paper, we introduce an utterance censorship system to check whether the chatbot's utterance is appropriate. If the speech is inappropriate, the censor will block it and give a negative reward to "punish" the chatbot. The censorship system is based on a character-level bidirectional LSTM model, and the chatbot receiving the reward from the censorship system "forgets" the learned offensive utterances. Experimental results show that our proposed architecture enables online learning chatbots to self-purify and that character-level LSTM is more appropriate for the utterance censorship task compared with classical word-level LSTM model.