Securing Automatic Speaker Verification Systems Using Residual Networks
Nidhi Chakravarty, Mohit Dua · Advances in information security, privacy, and ethics book series · 2024
Spoofing attacks are a major risk for automatic speaker verification systems, which are becoming more widespread. Adequate countermeasures are necessary since attacks like replay, synthetic, and deepfake attacks, are difficult to identify. Technologies that can identify audio-level attacks must be developed in order to address this issue. In this chapter, the authors have proposed combination of different spectrogram-based techniques with Residual Networks34 (ResNet34) for securing the automatic speaker verification (ASV) systems. The methodology uses Mel frequency scale-based Mel-spectrogram (MS), gamma scale-based gammatone spectrogram (GS), Mel filter bank-based Mel frequency cepstral spectrograms (MCS), acoustic pattern-based acoustic pattern spectrogram (APS), gammatone filter bank-based gammatone cepstral spectrogram (GCS), and short-time Fourier transform-based short Fourier spectrogram (SFS) methods, one by one, at front of the proposed audio spoof detection system. These spectrograms are individually fed to ResNet34 for classification at the backend.