Audio Deepfake Detection using Deep Scattering Network

Utsav Avaiya, Naman Badlani, Rajas Baadkar, Kiran Talele · 2024

The rise of audio Deepfakes has raised serious concerns about the authenticity of audio content, affecting media, politics, and cybersecurity by enabling misinformation, fraud, and identity theft. To address these challenges, we propose a novel approach to Deepfake detection using a Deep Scattering Network (DSN). Unlike traditional methods that rely on Convolutional Neural Networks (CNNs), which require massive amounts of training data, our approach leverages the Scattering Transform to capture multi-scale time-frequency features, enhancing the detection of subtle anomalies in manipulated audio while requiring significantly less data for training. Additionally, DSNs employ pre-defined multi-scale wavelet filters rather than data-driven linear filters, offering a more interpretable model while retaining a hierarchical structure. Evaluation of a publicly available dataset shows that the DSN outperforms the standard CNN-based models, particularly in scenarios with low-quality or heavily distorted audio. This research underscores the potential of scattering networks in bolstering the security and reliability of audio authentication systems.

Read the paper · More papers on PaperTik