Spectrogram-Based Analysis and Detection of Deepfake Audio Using Enhanced DCGANs for Secure Content Distribution

Jovelin M. Lapates, Bobby D. Gerardo, Ruji P. Medina · 2024

While DCGAN as deep learning model utilizing spectrogram, allows for detection of deepfake audio, it is prone to overfitting which affects its ability to discriminate between real and fake audio. In this study, batch normalization is incorporated into both the generator and discriminator to address training instability. The datasets, consisting of real human speech and DeepFake renditions generated through Retrieval-based Voice Conversion (RVC), are categorized into ‘REAL’ and ‘FAKE’ (FoR) classes and preprocessed using Audacity and Sonic Visualizer. The paper introduces an enhanced DCGAN model for augmenting samples in voiceprint recognition and evaluates various spectrogram techniques—Mel-Spectrograms, GTCC, MFCC, and Chroma-CQT—to improve detection accuracy. The model achieved a training accuracy of 92.86% and a validation accuracy of 91.67%, underscoring the potential of advanced deep learning methods to ensure audio authenticity against deepfake threats.

Read the paper · More papers on PaperTik