One-Class Fake Speech Detection Based on Improved Support Vector Data Description

Jinghong Zhang, Xiaowei Yi, Xianfeng Zhao · Security and Communication Networks · 2023

With the development of deep neural synthesis methods, speech forgery techniques based on text-to-speech (TTS) and voice conversion (VC) pose a serious threat to auto speaker verification (ASV) systems. Some studies show that the attack success rate of deep synthetic speech on ASV systems can reach about 90%. Existing detection methods improve the detection generalization of known forgery methods by a lot of training data, but the detection effect and robustness against unknown methods are poor. We propose an anti-spoofing scheme based on one-class classification for detecting unknown synthetic. We implement deep support data description to capture the feature of bonafide speech. An autoencoder structure is introduced to enhance the detection performance. The proposed method is only trained on native speech, which reduces reliance on large amounts of fake speech. Our method achieves an equal error rate of 8.10% on the evaluation set of ASVspoof 2019 challenge and outperforms other state-of-the-art methods. In the generalization test, the proposed method can reach the equal error rate of 15% on “In-the-wild” dataset and 23% on FoR dataset, which is lower than that of other advanced algorithms.

Read the paper · More papers on PaperTik