MADD: A Multi-Lingual Multi-Speaker Audio Deepfake Detection Dataset

Xiaoke Qi, Hao Gu, Jiangyan Yi, Jianhua Tao, Yong Ren, Jiayi He, Siding Zeng · 2024

AI-driven advancements in speech synthesis and voice conversion, now are able to convincingly emulate human speech, have made a growing challenge for investigators and the judicial system to discern between genuine and artificially generated audio. Creating an effective audio deepfake detector necessitates a large-scale and high-quality data. Existing datasets focus on monolingual data for high-resource languages. In this paper, we have constructed a multi-lingual multi-speaker audio deep-fake dataset, named MADD. The sources of MADD are de-rived from Common Voice and Gigaspeech2. Leveraging var-ious deep speech synthesis and voice conversion technologies across 6 languages, the MADD dataset comprises a collection of 129,990 deep synthesis utterances, with a total duration of 155.66 hours from 288 speakers.

Read the paper · More papers on PaperTik