Residual-based feature enhancement for forgery audio detection

Wei Zheng, Xia Ling, Han Hai · 2023

In recent years, speech synthesis technology has become increasingly advanced, leading to a proliferation of forged audio content on the internet, which poses significant threat to individuals and society. Many studies have utilized a range of deep learning-based techniques to differentiate fake audio content, but the features used in these studies are often limited in their rich and generalizable characteristics. In this paper, we propose a novel fake voice detection technology that utilizes the wav2vec2 model for feature extraction along with a custom-designed residual-based detection module to augment the detection of fake audio content with greater accuracy and precision. Additionally, we incorporate a data augmentation method to improve the performance of the model and enhance its ability to generalize. We trained our model on the ASVspoof2019 dataset and evaluated it on the LA and DF datasets of the ASVspoof2021 dataset. Supplementary experiments demonstrated that our approach achieved state-of-the-art detection performance and illustrated its effectiveness and applicability.

Read the paper · More papers on PaperTik