Detecting Speech Deepfakes through Improved Speech Features and Cost Functions

Oscal Tzyh-Chiang Chen, Yun-Chia Hsu, Tun-Sheng Yang · 2024

Due to the rapid development of artificial intelligence, deepfakes have given rise to malicious individuals exploiting voice manipulation as a means to deceive unsuspecting victims. In this study, we address the crucial task of defining acoustic feature parameters suited for capturing both breathing sounds and low-frequency burst characteristics, achieved through the meticulous configuration of the constant Q transform. The proposed model is constructed upon the foundation of a multi-group channel-wise gated Res2Net-50 architecture, complemented by the incorporation of three refined highway units equipped with residual connections. To effectively separate the distributions of positive and negative samples within the feature space, we’ve ingeniously tailored the cost function of one-class learning. Furthermore, this study introduces an improved cost function by synergizing one-class learning, supervised contrastive learning, and modified one-class learning with learnable ratios. By using the ASVspoof 2019 LA dataset, the proposed model achieves an Equal Error Rate (EER) of 0.6 8% and min t-DCF of 0.0194 in testing, surpassing the conventional models within the single system category. These results highlight the superior performance of the proposed model in the realm of voice deception detection.

Read the paper · More papers on PaperTik