Audio Deepfake Detection and Classification

B. Sarada, TVS. Laxmi Sudha, Meghana Domakonda, B. Vasantha · 2024

AI Voice Cloning, also called audio deepfakes is a highly advanced process that utilizes Artificial Intelligence to create a replica of a human voice. There is no doubt that this technology has revolutionized the way we interact with machines and has immense potential for various industries. This technology is used to create new identities or to steal the identities of the original voice owner and spread misinformation using cloned audio. This paper aims to differentiate between cloned voice and original voice using GAN and random forest. A generative adversarial network (GAN) is a deep learning architecture where two neural networks engage in a competitive dynamic within a zero-sum framework, striving to enhance the precision of these predictions. The Synthesized Data contains a lot of disturbances in the background which are generally referred to as Noise. To decrease this noise from the speech signals Spectral Subtraction is used. Feature extraction is done through zero crossing and a Random Forest classifier is used. By this classification, 100% accuracy has been acquired and other metrics such as precision, recall, and F1 score are also approximately equal to 100%. For analysis, a folder of 88 audio is considered.

Read the paper · More papers on PaperTik