Towards Secure Audio: Deepfake Detection with CNN and LSTM Networks

Tushar Bhagat, Neha Borge · INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2025

In recent years, advancements in artificial intelligence have led to a surge in the generation of synthetic and manipulated audio, commonly referred to as "deepfake audio." While these technologies offer advantages across various domains, they also present serious security and ethical concerns, particularly in contexts where the authenticity of audio is critical. This paper introduces a novel deep learning-based approach for detecting deepfake audio using a combination of Convolutional Neural Networks (CNN), Long Short-Term Memory (LSTM) networks, and an Attention mechanism. The proposed architecture utilizes CNNs to extract high-level spatial features from audio spectrograms, while the LSTM network captures the temporal dependencies inherent in audio sequences. The integration of the Attention mechanism further enhances the model's ability to focus on key segments of the audio that are more likely to contain deceptive artifacts. Through comprehensive experimentation on publicly available datasets, our model demonstrates superior performance in terms of accuracy and robustness compared to traditional and standalone deep learning models. These findings underscore the potential of hybrid architectures in effectively addressing the challenges of deepfake audio detection and contribute to the development of trustworthy audio verification systems. KEYWORDS deepfake audio detection, synthetic audio, machine learning, digital forensics, neural networks, feature extraction, deep learning, audio synthesis, data integrity, security

Read the paper · More papers on PaperTik