Dialogue-Level Data Augmentation for Conversation Derailment Forecasting and Topic-Shift Detection
Nerses Yuzbashyan, Nikolay Banar, Walter M. P. Daelemans · 2024
Counteracting antisocial behavior on-line becomes an increasingly challenging task, as social media platforms continue to gain popularity.In addition, it often requires predicting in advance when an online conversation is heading towards derailment in order to take actions before any harm is done.Developing systems for such purposes requires a large amount of labeled data, which is difficult to collect and extremely expensive to annotate manually.In such conditions, data augmentation could be an attractive alternative.In this paper, we address the conversation derailment detection and forecasting tasks and conversation topic-shift detection task.These tasks require augmentation at the dialogue level, which presents unique challenges compared to other data augmentation approaches.We propose three methods for generating synthetic dialogues using large language models (LLMs) to augment training datasets without the need for additional data collection.Our results demonstrate that while the proposed methods yield improvements for scarce data, they cannot overcome the ceiling effect when data is abundant.