Improving Fairness in Synthetic Speech Detectors
Amit Kumar Singh Yadav, Kratika Bhagtani, Paolo Bestagini, Edward J. Delp · 2024
Many methods have been proposed which can effectively detect synthetic speech. However, a recent study demonstrates that they exhibit bias and a higher false positive rate for bona fide speech from speakers with stuttering speech-impairment as compared to fluent speakers. This limits deploy-ment of these detectors, as this bias can have significant societal and political consequences and can erode the reputation of social platforms using such detectors. This bias may have arisen from the bias in training data used for these detectors. Creating synthetic and bona fide speech with stuttering, and adding it to the training set for mitigating bias, can be time-consuming. In this work, we propose StutterAug, a set of augmentations that simulate three major types of stuttering in speech, namely repetition, prolongation and blocks. We test StutterAug on 3 synthetic speech detectors and examine bias on stuttering speech using more than 28K bona fide stuttering speech. Our results show that detectors trained with StutterAug have on average 13% less bias relative to detector trained without StutterAug. StutterAug also leads to an average relative improvement of 27.96% in detection performance on ASVspoof2019 dataset and 11.27% in generalization performance on In-the-Wild dataset compared to baseline detectors trained without StutterAug.