Detection of Audio Attacks (Deepfake) Using Time-Based and Cepstral Domain Features with Stacking Classifier

Zefanya Darma Putri, Vera Suryani · 2025

Audio deepfake, or the manipulation of sound, imitates or changes the original voice and can be used for fraud and defamation. The purpose of this study is to improve the accuracy of deepfake audio detection by using a stacking classifier with the best parameters of SVM, Random Forest and Logistic Regression as base learners. This research focuses on exploring the combination of features from various domains to enhance deepfake audio detection, used six types of features, such as Mel-Frequency Cepstral Coefficients (MFCC), Spectral Rolloff, Spectral Contrast, Bandwidth, Zero-Crossing Rate (ZCR) and Root Mean Square (RMS). The author used The Fake or Real dataset, which was created using a text-to-speech model and divided into four sub-datasets: for-rerec, for-2-sec, for-norm, and for-original. Additionally, a machine learning-based stacking classifier is implemented and compared with conventional machine learning and deep learning methods to evaluate its effectiveness. This approach aims to address the limitations of previous models, which struggle to handle variations in synthesized speech within deepfake datasets. The experimental result of this system has 98-99% accuracy of testing and 97-99% accuracy of validation. This research demonstrates the effectiveness of the stacking classifier approach in detecting real or fake audio deepfakes and shows improvement over previous research.

Read the paper · More papers on PaperTik