Teager Energy Cepstral Coefficients for Audio Deepfake Detection
Ritik Mahyavanshi, Chaitanya Reddy, Arth J. Shah, Hemant A. Patil · 2024
Audio deepfakes have emerged as a significant concern, as these convincingly mimicked synthetic audio and can be exploited to manipulate targeted groups by altering speech information. No prior research has explored the use of Teager Energy Operator (TEO)-based features for ADD task. This paper discusses a novel technique for distinguishing deepfake audio from real ones using Teager Energy Cepstral Coefficients (TECC) features, which captures the energy fluctuations of audio signals, making it a promising approach for detecting real vs. deepfake audio. For classification, the ResNet-50 model is employed in combination with the FoR (Fake or Real) dataset. Experimental results demonstrate the efficacy of this method, achieving an Equal Error Rate (EER) of approximately 7.53 % and an accuracy of 92.55 % on static TECC feature sets. The findings indicate that TECC features outperform traditional techniques, such as Mel Frequency Cepstral Coefficients (MFCC), and Linear Frequency Cepstral Coefficients (LFCC). This superior performance highlights the significant potential of TECC features for further research in deepfake detection.