Deepfake Audio Detection Using Feature-Based and Deep Learning Approaches: ANN vs ResNet50

Reham Mohamed Abdulhamied, Sarah Naiem, Mona Mohamed Nasr, Farid Ali Moussa · International Journal of Advanced Computer Science and Applications · 2025

The proliferation of algorithms and commercial tools for generating synthetic audio has sparked a surge in mis- information, especially on social media platforms. Consequently, significant attention has been devoted to detect such misleading content in recent years. However, effectively addressing this challenge remains elusive, given the increasing naturalness of fake audio. This study introduces a model designed to distinguish between natural and fake audio, employing a two-stage approach: an audio preparation phase involving raw audio manipulation, followed by modeling using two distinct models. The first model employed feature extraction through wavelet transformation, followed by classification using a machine learning Artificial Neural Network. The second model utilized ResNet50 architecture, a type of deep learning model, which resulted in improved accuracy. These findings underscore the effectiveness of deep learning approaches in audio classification tasks. Training data for the model is sourced from the DEEP-VOICE dataset, which comprises both genuine and synthetic audio generated by various deep-fake algorithms. The model’s performance is assessed using diverse metrics such as accuracy, F1 score, precision and recall. Results indicate successful classification of audio in 86% of cases. This research contributes to the field of Automatic Speech Recognition (ASR) by integrating advanced preprocessing techniques with robust model architectures to identify manipulated speech.

Read the paper · More papers on PaperTik