Leveraging Neural Vocoder Artifacts for Improved Synthetic Speech Detection

Jingxi Xue, Sanshuai Cui, Weinan Zhang · 2024

With the advancement of deep learning, speech synthesis technology has improved significantly, but this progress also increases the risk of speech forgery.In this paper, we enhance synthetic speech detection by fully leveraging neural vocoder artifacts and improving feature representation.Our approach enhances the model's ability to detect synthetic speech by identifying artifacts produced by neural vocoder.We introduce the Enhanced Feature Extraction Module and the Dynamic Layer Attention Module to improve the model's feature representation and enhance its sensitivity to artifact features, significantly boosting the performance of synthetic speech detection.Additionally, we apply the exponential moving average (EMA) during training to further stabilize the model.Our experiments show that this method achieves an EER of 0.03% on the public datasets LibriSeVoc and 3.10% on the ASVspoof 2019 LA task.Compared to the LFCC-LCNN baseline, these results represent relative reductions of 78.6% and 73.3%, respectively.

Read the paper · More papers on PaperTik