Disfluency Detection for Real-World Scenarios
Jianbang Ding, Suiyun Zhang, Dandan Tu · 2024
Most existing methods for disfluency detection mainly rely on human-annotated data, which can be costly to obtain in real-world scenarios. Previous work showed that the benefit of data augmentation approaches to alleviate this problem. However, these studies mainly focused on the simple types of synthetic disfluencies that do not match the natural speech. In this work, we propose a novel end-to-end data augmentation technique that prompts large language models to generate natural and diverse disfluent texts from real examples. We further use this augmented data for pretraining and leverage it for the task of disfluency detection. We also propose a deletion-checking method to prevent wrong deletions hence better adapting the detection model to real-world scenarios. We conduct extensive experiments on the publicly released corpus, Switchboard, as well as our proprietary Chinese dataset. Experimental results show that our approach significantly outperforms previous baselines and achieves state-of-the-art performance (94.3 F-score) on English Switchboard corpus. Further qualitative analysis and an ablation study provide more insights into our approach.