Dual Acoustic Feature Fusion for Enhanced Audio Deepfake Detection Using VGG-16 Architecture: Mitigating Speech Tampering with MFCC and ELTP

Candra Ahmadi, Siao-He Wang, Shih-Ping Chiu, Jiann-Liang Chen · 2024

In recent years, the rise of audio deepfakes, where artificial intelligence (AI) is used to manipulate or synthesize speech, has posed significant challenges in distinguishing between authentic and falsified audio content. Despite advancements in detection methods, current systems often struggle with accurately identifying deepfakes due to the complexity of modern voice cloning technologies. This study addresses this gap by proposing a novel dual-input convolutional neural network (CNN) model utilizing Mel-Frequency Cepstral Coefficients (MFCC) and Extended Local Ternary Patterns (ELTP) for enhanced audio forgery detection. Here, we present a system based on the VGG-16 architecture, which integrates MFCC and ELTP features to detect manipulated audio with high accuracy. Our approach achieves a detection accuracy of 94.21% on the Fake-or-Real and ASVspoof 2019 datasets, demonstrating a robust improvement over existing single-feature methods. These results suggest that combining MFCC and ELTP features can better capture both the spectral and textural properties of audio, leading to more reliable detection of audio tampering. This advancement has important implications for cybersecurity and the integrity of voice-based authentication systems, providing a stronger defense against potential fraud and misinformation through audio deepfakes.

Read the paper · More papers on PaperTik