Analysis of Amplitude and Frequency Perturbation in the Voice for Fake Audio Detection
Kai Li, Yao Wang, Le-Minh Nguyen, Masato Akagi, Masashi Unoki · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022
Fake audio detection (FAD) aims to detect fake speech generated by advanced voice conversion and text-to-speech technologies. Recently, the quality of synthesized speech has significantly improved due to the remarkable development of deep neural networks. However, it is still easy for humans to identify fake speech by perceiving pathological prosody in a voice. Pathological prosody is significantly related to the amplitude and frequency perturbation (AFP) in the voice and provides essential cues to identify fake speech. This paper proposed to analyze AFP differences in the voice using the jitter and shimmer features. According to the statistical analysis of AFP features, the continuous-shimmer feature (CS3) can effectively separate genuine and fake speech signals. Moreover, static and dynamic CS3 features were combined with a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system, and experiments on datasets of the Audio Deep Synthesis Detection Challenge (ADD2022) were carried out. The results of the experiments show that both the static and dynamic shimmer features of voice can provide complementary knowledge to the traditional spectrum-based FAD systems.