Enhancing Synthesized Speech Detection Using Modified Group Delay and Self-Supervised Compact Convolutional Transformer
Soulaiman Moualla, Oumayma Al Dakkak, Omar Hamdoun, Assef Jafar · IEEE Access · 2026
The escalating sophistication of synthesized speech necessitates detection systems that target fundamental physiological constraints rather than surface artifacts. Our approach leverages Modified Group Delay (MGD) to capture the vocal tract characteristics that remain challenging for synthesizers to replicate. We then use this robust representation to train a Compact Convolutional Transformer (CCT) model via self-supervised learning. This training is based on masked prediction, inspired by models like Wav2Vec2, using an Information Noise-Contrastive Estimation (InfoNCE) loss to force the model to learn the local, deep, and contextual structure of real human speech. The trained CCT thus becomes a powerful encoder, producing discriminative embeddings. Finally, a simple supervised binary classifier is trained on top of these embeddings to distinguish between real and synthesized speech, creating a hybrid detection system resilient to novel and evolving synthetic threats. This method achieved state-of-the-art performance when evaluated on the ASVspoof 2019 LA dataset, with an Equal Error Rate (EER) of 0.0093%, thus performing a new benchmark without augmentation. Furthermore, it demonstrated superior generalization, with an EER of 0.61% on the more challenging ASVspoof 2021 LA dataset and excellent performance with an EER of 6.13% on the In-The-Wild dataset.