Significance of Lower Frequency Regions for Audio Deepfake Detection

Arth J. Shah, Hemant A. Patil · 2024

Deepfake audios are created using deep learning methods. Audio security, due to deepfake audio creation, has been a severe issue in recent days. Many techniques have been explored to detect attacks by fake audio generators. Similar to Spoofed Speech Detection (SSD), audio deepfake systems also rely on characteristics of live speech to detect whether the audio is deepfake or real, however, they aim to fool humans instead of voice biometrics system. As the attackers are free to mount various types of audio attacks, Audio Deepfake Detection (ADD) plays a vital role in defending against fake AI/deep learning-generated audio attacks. For this study, we explored various lower frequency-based acoustic features, such as Generalized Morse Wavelet (GMW)-based features in combination with Mel spectrogram-based features, using Convolutional Neural Network (CNN)-based classifier for ADD task. For performance comparison, we employed the traditional spectrogram. Further, feature-level data fusion (instead of score-level) of proposed GMW-based with Mel spectrogram feature gave 1.93% improvement in overall accuracy.

Read the paper · More papers on PaperTik