Deepfake Audio Detection Using Generative Adversarial Network-Based Spatial–Temporal Feature Learning

Vivek Ranjan, Dayashankar Singh · International Journal of Software Engineering and Knowledge Engineering · 2025

Deepfake audio detection is imperative in the digital era as it has a detrimental effect on individuals, organizations and countries. Deepfake audio is used more and more by cybercriminals for impersonation, fraud and voice phishing, necessitating further research on detection. Previous research in the area has largely used supervised learning, which provides effective results using labeled data for training. This paper proposes a supervised learning framework for detecting deepfake audio using a generative adversarial network (GAN) with spatial–temporal feature learning. The audio is first de-noised using a Wiener filter, and then key pitch and speech features, including Mel-frequency cepstral coefficients (MFCCs), pitch entropy, harmonic-to-noise ratio (HNR) and glottal flow, are extracted. These features are processed by a one-dimensional dilated spatial–temporal attention-based bidirectional gated recurring units (ODDST-CBG) to capture nuanced patterns. A Wasserstein GAN (W-GAN) discriminator is used for final classification. Experimental results on a large-scale fake versus real speech dataset demonstrate high performance, achieving 99.98% accuracy, 99.08% precision, 97.08% sensitivity and a 97.08% F1-score, surpassing existing baseline methods.

Read the paper · More papers on PaperTik