Deepfake Urdu Audio Detection using Spectral Features for Automatic Speaker Verification
Zarreen Sajjad, Hassan Alam, Ahmed Nouman, Areeba Shahzad Raja, Junaid Mir, Furqan Shaukat · 2024
Deepfake audio detection for automatic speaker verification in Urdu is a significantly underexplored area. This paper presents a deep learning-based deepfake audio detection method for low-resource language Urdu. As the Urdu language has different phonetics than English, the efficacy of spectral features extracted from 1D and 2D audio signal representations is assessed for spoofed Urdu audio detection. Towards this, four Deep Convolutional Neural Networks (DCNNs) are trained to learn spectral features from 2D spectrogram and scalogram images generated from a recently released specialized Urdu deepfake audio dataset. Also, a support vector machine classifier is trained on Mel-frequency cepstral coefficients extracted from 1D audio signals. Results reveal that the EfficentNetV2-B0-based DCNN classification model trained on spectrograms surpasses the tested models in deep fake audio detection for the Urdu language.