ASSD: An AI-Synthesized Speech Detection Scheme Using Whisper Feature and Types Classification

Chang Liu, Xiaolong Xu, Fu Xiao · IEEE Transactions on Audio Speech and Language Processing · 2025

Current efforts in AI-synthesized speech detection often suffer from decreased performance due to detectors overfitting to irrelevant features or fitting features specific to certain synthesis methods. This issue becomes severe when facing unseen synthesis methods or real-world scenarios. To address these challenges, we first construct a multi-scenario dataset named PolyFake. Subsequently, we propose an AI-Synthesized Speech Detection Scheme using Whisper Feature and Types Classification (ASSD). In the feature processing stage, this method adopts a dual-stream network. Firstly, it utilizes the proposed SC_Encoder to separate and reconstruct redundant features, thereby reducing attention on irrelevant features. Meanwhile, it extracts features through a pre-trained Whisper_Encoder to augment feature generalization ability. Then, features extracted by the two encoders are decomposed, concatenated, and fused in both time and frequency domains, concurrently constructing time and frequency domain graphs. Through heterogeneous stacking graph attention, temporal and spectral information is fused to obtain generalized features. In the detection stage, employing a multi-task learning strategy, we introduce a multi-classification task for predicting synthesis method types to further segregate generalized features, obtaining common features of different synthesis methods for synthesized speech detection. The proposed method achieves lower equal error rate compared to state-of-the-art methods, with a relative improvement of 27% in generalization performance under a 2% increase in parameter count, on the In_the_wild and ASVspoof 2021 DF dataset and the testing scenarios of PolyFake.

Read the paper · More papers on PaperTik