Cloud Voice Security: Anti-Deepfake via Graph Attention Aggregation and Post-Quantum Cryptography for Cloud Service
Weijiang Xia, Haipeng Peng, Lixiang Li, Yuning Qi, Yeqing Ren · IEEE Internet of Things Journal · 2025
With the development of content generative large models, spoofed speech can be synthesized more easily. However, the generalization of current cloud voice anti-spoofing detection model is insufficient, especially when facing attacks from unknown speech synthesis algorithms. Meanwhile, voice data is also susceptible to attacks during transmission. Therefore, we design an end-to-end voice anti-spoofing scheme, which can be applied in IoT cloud service. The scheme consists of an encryption transmission module and an anti-spoofing model. The encryption module is designed with post-quantum cryptography and chaos to protect transmission security. In the modle, we propose a new higher-order two-dimensional attentive statistics pooling (H2D-ASP) module to extract and aggregate more attention representations in spectral domain and temporal domain; And we propose a new channel-dependent self attention based graph aggregation (CSA-GA) module, which squeezes and aggregates spectral graphs and temporal graphs. Finally, we conduct experiments on the ASVspoof 5 Challenge deepfake database under the closed condition. The experiments show that the model in our scheme is a better single model which performs minDCF 18.28% better than the baseline model on the evaluation set. Without data augmentation, the model achieves minDCF of 0.2994, 0.3604, 0.581 on the development set, evaluationprog set and evaluation set, respectively. The last two proposed modules improve the vanilla model by 29.75% on evaluationprog set and 18.97% on evaluation set.