Attribution of Diffusion Based Deepfake Speech Generators
Kratika Bhagtani, Amit Kumar Singh Yadav, Paolo Bestagini, Edward J. Delp · 2024
Several text-to-speech (TTS) generation methods have been recently proposed which use diffusion models. Synthetic speech has been maliciously used for impersonation and to spread misinformation. Therefore, synthetic speech detection and attribution methods have been developed. Synthetic speech detection methods can detect synthetic speech. Synthetic speech attribution methods can identify the generator that was used for synthesizing a given speech. Existing attribution methods attribute a given speech to conventional speech generators, and their performance is demonstrated mainly on the ASVspoof2019 Dataset. In this work, we explore the attribution of latest, diffusion model based speech generators including commercial speech generation software. We experiment with four synthetic speech attribution methods, two of which demonstrate more than 99% attribution accuracy. These methods can also identify unseen speech generators which have not been used for their training.