Compressed Domain Invariant Adversarial Representation Learning for Robust Audio Deepfake Detection

Chengsheng Yuan, Yifei Chen, Zhili Zhou, Zhihua Xia, Yongfeng Huang · IEEE Signal Processing Letters · 2025

The primary aim of audio deepfake detection (ADD) is to thwart deception arising from forged audio generated through text-to-speech or voice conversion technologies. However, encoding speech signals using diverse compression algorithms introduces discrepancies that significantly impair the performance of existing countermeasure systems. To tackle these challenges, this letter proposes a robust audio deepfake detection method based on Compressed Domain Invariant Adversarial Representation Learning with Adaptive Token Pooling (DANet-ATP). This framework incorporates a Compression Codecs Discriminator (CCD) that, through adversarial learning in tandem with the backbone network, enhances the model's ability to extract more robust features across diverse compression codecs. Moreover, to efficiently prune redundant frame-level features while retaining vital spoofing cues, the letter designs a plug-and-play, parameter-free Adaptive Token Pooling module, significantly improving detection performance. Experimental results on the ASVspoof2021 DF dataset showcase the exceptional performance of the proposed model. Furthermore, a series of ablation experiments validate the validity and effectiveness of the proposed method.

Read the paper · More papers on PaperTik