Performance Evaluation of Enhanced DCGAN s for Detecting Deepfake Audio Across Selected FoR Datasets

Jovelin M. Lapates, Bobby D. Gerardo, Ruji P. Medina · 2024

As the generation of deepfake audio becomes more advanced, the need for effective detection techniques has grown significantly. This study explores the application of an enhanced DCGAN for detecting deepfake audio, with a focus on the Fake or Real (FoR) dataset. The model utilizes convolutional layers to process spectrograms derived from audio signals. By incorporating batch normalization in both the generator and discriminator, the model addresses training instability and improves convergence. The audio data, including real human speech and deepfake renditions generated through Retrieval-based Voice Conversion (RVC), were preprocessed using Audacity and Sonic Visualizer. The enhanced DCGAN was then trained on these spectrograms, leveraging adversarial training to improve detection capabilities. Performance evaluation of the model involved key metrics such as accuracy, precision, recall, and F1-score. The results revealed an accuracy rate of 98%, marking a 6.96% improvement over the standard DCGAN. Furthermore, the enhanced DCGAN exhibited superior precision and recall in distinguishing between real and fake audio samples.

Read the paper · More papers on PaperTik