Improving Imbalanced Dialogue Act Classification Using Cost-Sensitive Learning
Takaaki Miyagi, Satoshi Endo · 2022 Joint 12th International Conference on Soft Computing and Intelligent Systems and 23rd International Symposium on Advanced Intelligent Systems (SCIS&ISIS) · 2022
Dialogue acts represent the content or intention of the speaker’s utterance, and are classified into different types, which are called dialogue act labels. Since dialogue acts are helpful in understanding the dialogue content and generating sentences, many studies have been conducted on the classification of dialogue acts. In particular, the dialogue act classification of a response is a problem in predicting dialogue act labels of response sentences and is effective for generating response sentences if the labels are correctly classified. However, the dataset used in this study is unbalanced, with the utter and understand labels accounting for more than 70 % of the nine dialogue act labels. Therefore, the majority labels are predicted more often and the minority labels are not predicted at all. This paper aims to solve the problem of minority labels not being predicted. Increasing the loss value when minority labels are misclassified, lets the model learn to capture the features of minority labels. We believe that cost-sensitive learning is effective for classifying imbalanced datasets because it allows us to adjust the loss value in case of misclassification. The cost-sensitive learning used in this experiment can be roughly divided into two types depending on the setting of the cost value. The first is a method that sets a cost value for each class, and the second is a method that sets a higher cost value when a minority label is misclassified as a majority label. Using these methods, we have improved the Recall and F1 compared to the conventional methods and reduced the misclassification rate of minority labels as majority labels. We confirmed that the cost-sensitive learning method is effective for imbalanced datasets. However, the cost-sensitive learning method did not improve accuracy.