Audio deep fake detection using LSTM-RNN

Maidhili Mohan K, Karthika K Balachandran, Rameela Ravindran K., Risma P.M. · INTERNATIONAL JOURNAL OF CURRENT SCIENCE · 2025

The need for digital content authentication has arisen from the emergence of deep fake technology. The development of generative models made it simple to create and edit digital content including images, audio clips, videos, etc. Malicious uses for audio deep fakes include phony voice calls, impersonation and the fabrication of audio evidence. In such scenario, developing website to identify deep fake audios becomes essential. The proposed method makes use of deep neural networks to distinguish between real and fraudulent audios. Users have access to a website where they can submit audio files for evaluation. The process of detection is carried out by the spectral analysis of the audio file. The overall goal of the paper is to identify and prevent the spread of altered digital audios or audio deep fakes. Additionally, the paper enhances critical thinking and mature content consumption, emphasizing the user awareness and literacy in regulating the spread of audio deep fakes. The study mainly focuses on the effect of the Long Short-Term Memory (LSTM) and the Recurrent Neural Networks (RNN) on detection of the deep fake audio to accomplish the mission of the paper. Temporal dependencies in audio data are taken consideration using LSTM-RNN architectures, which allow for the detection of minor patterns that may indicate tampering. The study makes use of fororiginal dataset from Fake-or-Real (FoR) dataset. The Fake-or-Real dataset is divided into four datasets: forrece, for-2-sec, for-norm, and fororiginal, where the for-original dataset is the collection of other three datasets without any pre-processing and this work uses the original dataset. Mel Frequency Cepstral Coefficients (MFCC) are the primary feature extraction method employed in this study. Overall, the study emphasizes the need of advanced technology and user education to solve the problems caused by deepfake audio in the current digital environment.

Read the paper · More papers on PaperTik