Advanced Sequential Modeling for DeepFake Audio Identification

Priyadarshan S. Dhabe, Nitin Choudhary, Ayush Vidhale, Yash Munde, Muaz Sayyed, Netal Zanwar · International Journal for Research in Applied Science and Engineering Technology · 2024

Abstract: The research paper is about audio DeepFakes, which are fake audio clips that sound just like real people’s voices. These can be used to spread false information or pretend to be someone else, which is a big problem.The paper looks at different ways to tell if an audio clip is a DeepFake or not. It uses advanced computer techniques, like generative adversarial networks (GANs), to do this. The paper tests different methods, like using raw sound waves, Mel-frequency cepstral coefficients (MFCCs), and linear frequency cepstral coefficients (LFCCs) as inputs. One of the methods tested is the Time-Domain Synthetic Speech Detection (TSSD) model, which takes raw audio waveforms as input. The tests are done on the WaveFake dataset, which has synthetic audio made by six different GAN architectures. The paper measures how well each method works using things like the equal error rate (EER), F1 score, and area under the receiver operating characteristic curve (ROC AUC). The results show that some methods, like the shallow CNN and TSSD architectures, are good at detecting audio DeepFakes. But, the paper also points out that there’s still room for improvement, especially for cases where the fake audio is made to trick the detection methods.

Read the paper · More papers on PaperTik