Audio Integrity Verification: Detecting Manipulated and Fake Speech

Mr. K. Amrutasagar, D Keerthi, P. Nikhil, Prasanna N. Lakshmi, Varsha R Ajith · Industrial Engineering Journal · 2025

One important area of research is deepfake audio detection, which separates real human voices from speech that has been modified or produced artificially. Generative models like WaveNet, Voice Conversion, and Text-to-Speech (TTS) synthesis have greatly enhanced the quality and realism of deepfake audio due to the quick development of artificial intelligence. This has raised significant ethical and security issues in a number of domains,including media, cybersecurity, and forensic investigations. In order to analyze Mel-Spectrograms and MelFrequency Cepstral Coefficients (MFCCs), this study suggests a deepfake audio detection framework that makes use of Convolutional Neural Networks (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) networks. These attributes enable the model to accurately discriminate between synthetic and real audio by capturing the spectral and temporal aspects of speech. A balanced training environment is ensured by the 94,734 audio samples in the dataset utilized in this study, which is evenly distributed between actual and false recordings. To improve model performance, preprocessing methods such time-frequency domain analysis, feature scaling, and noise removal are used. Using demanding experimental settings, the suggested CNN-BiLSTM architecture is trained and assessed, obtaining 98% accuracy and proving its resilience in identifying deepfake speech. In order to prevent audio forgeries and improve the security of voice-based authentication systems, the results of this study emphasise the significance of hybrid deep learning architectures. In order to increase the scalability and versatility of deepfake detection models, future research will investigate the merging of self-supervised learning strategies with real-time detection methods.

Read the paper · More papers on PaperTik