Cross-Stage Mutual Distillation for Replay Attack Detection
Yinqing Cheng, Lasheng Zhao, Ling Wang, Han Wang · 2024
Automated speaker verification systems face a higher risk of replay attacks in practice. However, the existing studies face the problem of limited detection capabilities and insufficient use of shallow fine-grained information. To address these issues, we propose the cross-stage mutual distillation(CS-MD) framework, which involves two models learning from a deep network output of each other in different stages of training. This mutual learning approach enhances the ability of shallow networks to capture fine-grained speech information. Additionally, we use an attentional feature fusion module to integrate shallow information more effectively. The multi-scale attention mechanisms in this module can combine local and global speech features while preserving detailed information. Experimental results on the ASVspoof 2019 physical access dataset demonstrate that our proposed method outperforms state-of-the-art methods in terms of EER and min t-DCF metrics, validating the effectiveness of our CS-MD framework.