Div-Df: A Diverse Manipulation Deepfake Video Dataset

Deepak Dagar, Dinesh Kumar Vishwakarma · 2023

Recent advances in image and video manipulation have given rise to grave concerns. Deepfake technology employs deep learning techniques to produce astoundingly lifelike content. Deepfakes are risky since they have the ability to counterfeit someone’s identity by replacing their face with that of another person or generating random noise in the mouth area. Additionally, with just a few seconds of audio, AI-based deep learning models can replicate any person’s voice. Detecting such videos is the only promising defense against such fraudulent data. Several deepfake datasets have been made available to help in deepfake detector training and testing, including DF-TIMIT [1], FaceForensics++ [2], Celeb-DF [3], DFDC [4], Deeperforensics1.0 [5], etc. Even though this has significantly improved deepfake detection methods, they are still unable to capture real-world scenarios entirely, as most of the dataset is face-swap manipulation. To bridge this gap, we have proposed a Div-DF dataset containing various types of video manipulation like face swap, facial reenactment, and lip-sync. This dataset is composed of 150 real videos of different celebrities of different professions and 250 deepfake videos (100 face-swap videos, 100 facial reenactment videos, and 50 lip-sync videos). Deepfake videos are synthesized using state-of-the-art Face-Swap GAN(FSGAN) and the Wav2Lip method. The dataset contains high-quality samples of face-swapped and lip-sync videos, while the samples of face-re-enactment are of average quality. We have tested state-of-the-art detection and image classification models to standardize our dataset’s baseline evaluation of various detection methods. We have done a comprehensive assessment along different metrics and found that our dataset is challenging and represents real-world samples.

Read the paper · More papers on PaperTik