Mitigating Replay Spoofing Attacks in Speaker Verification for Secure and Trustworthy Voice Interactions
Mohammad Asgari, H. Hejazi · 2024
One of the main concerns of smart cities and systems is increasing their security. Audio fake replay is one of the attacks on speaker authentication (SV) systems. This article deals with the issue of reducing the recognition errors of audio replay attacks as one of the types of spoofs while increasing its generalizability by using a new model. This model is a combination of Convolutional Neural Network (CNN) and Vision Transformer (ViT) called CNN-ViT, which has the capabilities of both models in extracting local and global features. Using ASVSpoof 2017 challenge dataset and its augmentation reaches the equivalent error rate difference (EER) to 2.38, which is about 23% improvement in generalizability compared to the best case of the previous baseline methods.