Synthetic Speech Detection with Wav2vec 2.0 in Various Language Settings

Branimir Dropuljić, Miljenko Šuflaj, Andrej Jertec, Leo Obadić · 2024

Synthetic speech detection plays an important role in fending off ever-increasing malicious use of voice deepfake technologies. However, its robustness and generalization have not yet been explored in diverse language settings. In this paper, we primarily analyze how such a system is affected by: (i) biases caused by different textual domains within human and synthetic samples, (ii) unseen languages, and (iii) non-native speech. Two human speech datasets, FLEURS and ARCTIC (CMU and L2), were extended with generated text-to-speech (TTS) samples. The results indicate that the wav2vec 2.0 based models are agnostic to the aforementioned points.

Read the paper · More papers on PaperTik