An End-to-End Audio Transformer with Multi-student Knowledge Distillation algorithm for Deepfake Speech Detection

Weidong An, Ruwei Li, Haoyu Ge, Man Li, Huaiyu Li · 2024

An increased prevalence of fraudulent techniques has revealed the limitations in performance and detection speed of existing Spoofed Speech Detection(SSD) algorithms. To address these challenges, a more stable and rapid algorithm is proposed in this paper. Firstly, a novel feature extraction algorithm is introduced, in this algorithm we employing an end-to-end extraction frontend combined with a feature smoothing mechanism to extract more robust feature representations. Secondly, a one teacher and multi-student knowledge distillation system, guided by an Audio transformer as the teacher model. This system comprises two distinct networks: the teacher network and student network. Through a one-teacher and multiple-students knowledge distillation structure, the model achieves faster detection speeds without compromising performance, meeting the requirements for real-time processing. Finally, the algorithm utilizes the ASVspoof2021 LA dataset to simulate unknown attacks and employs pseudo labels generated by the teacher model to train the students model, thus enhancing the system's capability to handle increasingly variable unknown attacks in the future. Experimental results demonstrate that on the ASVspoof2019 evaluation set the proposed algorithm reaches optimal performance with the minimum model parameters that only 0.33M. Moreover, on the ASVspoof2021 LA and ASVspoof2021 DF evaluation sets, the algorithm proposed in this paper achieves performance close to the state-of-the-art (SOTA) algorithms while requires only 7.64ms for single speech inference on a CPU, fulfilling the real-time processing criteria.

Read the paper · More papers on PaperTik