Deep Learning based Spoof Detection: An Experimental Study

Nabeel Koya A, Joyanta Basu, Waquar Ahmad, P. V. Sudeep · 2023

This paper presents a novel approach to detect audio spoofing in the ASVspoof 2019 dataset. The proposed method utilizes Spectral Delta Coefficient (SDC) as a feature vector and employs Residual Neural Networks (ResNet) as a classifier. Audio spoofing poses a considerable challenge in forensic investigations and Automatic Speaker Verification (ASV) systems, with impersonation, voice conversion, speech synthesis, and replay attacks being common forms of spoofing techniques. Modern methods for generating spoofed speech using voice conversion and speech synthesis have reached a level of perceptual indistinguishability from genuine speech. Additionally, replay attacks allow perpetrators to deceive an ASV system by replaying a recorded voice of a legitimate human speaker. The paper evaluates and compares different feature vectors and classifier combinations. The feature vectors used in this study include Mel-frequency cepstral coefficients (MFCC), Constant Q Cepstral Coefficients (CQCC), Linear Frequency Cepstral Coefficients (LFCC), Spectrogram, and Spectral Delta Coefficients (SDC). Various classifiers, such as Neural Networks, Gaussian Mixture Models (GMM), Support Vector Machines (SVM), and Residual Networks (ResNet), are employed. Equal Error Rate(EER) is the evaluation metric used in the ASV spoof challenge. The results demonstrate that the SDC-ResNet combination outperforms other combinations, including LFCC-GMM, CQCC-GMM, MFCC-GMM, MFCC-SVM, SDC-GMM, SDC-SVM, and Spec-ResNet, in terms of EER.

Read the paper · More papers on PaperTik