Utilizing Multi-modal Emotion Information in Dialogue Strategy Classification
Jin Yea Jang, Ji-Eun Kim, Minyoung Jung, Hyedong Jung, Saim Shin · 2020
Dialogue models that automatically generate utterance sentences are being actively studied with the increase of data and the development of deep neural network technology. Research is also being conducted using information from multiple modalities to create dialogues that are more natural and reflect various aspects in the dialogue. In this paper, we focus dialogue strategy, which is important for understanding a natural dialogue flow. We implement a deep-learning based, dialogue strategy classifier based on the data labeled with emotion information from multiple modalities (i.e., text, sound, and image). Results of our experiment indicate that the use of emotion information from multi-modalities can help classify dialogue strategies with approximately 10% performance increase compared to the baseline. Results also show that the emotion information from the sound modality contributes to dialogue strategy classification more than that from the text and image modalities.