Audio Spoof Detection using Deep Residual Networks based Feature Extraction: Unveiling Synthetic, Replay and Mimicry Threats
Nidhi Chakravarty, Mohit Dua · 2024
With the increasing prevalence of artificial intelligence-driven audio systems, the vulnerability to adversarial attacks, specifically synthetic, replay and mimicry, has become a critical concern. This paper has proposed a deep Residual Network as a feature extractor to resolve this problem. This proposed approach is divided into two phases: Frontend and Backend. During the initial phase, the audio data underwent Mel spectrogram generation, and a modified ResNet50 architecture has been employed for feature extraction from the resulting Mel spectrograms. Subsequently, Linear Discriminant Analysis (LDA) has been applied as a dimensionality reduction technique to optimize the complexity of the extracted features. In the second phase, the features selected by LDA have been utilized to train machine learning (ML) algorithms, including Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbour (KNN), and Naive Bayes (NB). The training process involved utilizing the ASVspoof 2019 Logical Access (LA) and Physical Access (PA) training partitions, while the proposed model's validation and evaluation have been conducted using the development and evaluation partitions of the same dataset. Also, the Voice Impersonation Corpus in Hindi Language (VIHL) dataset has been used to detect mimicry attacks. The proposed model has achieved an EER of 2.4%, 2.2%, and 4.1% for synthetic, replay and mimicry audio attack detection, respectively.