Decoding Voice Authenticity: An Exploration of Deep Learning and Audio Feature Techniques

Swathi Tejah Yalla, M. Raju, D. Nagaraju, Mohammad Farhaan Ali · 2025

The increasing sophistication of spoofing, mimicry, and deepfake technologies exposes critical vulnerabilities in voice authentication systems, including the inability to generalize across diverse attack types, reliance on narrow feature sets, and a lack of interpretability, which seriously limits their applicability in the real world. This review will analyze the state-of-the-art techniques for Logical Access, Presentation Attacks, and Deepfake manipulations. It points out the limitations of the traditional models and progress based on modern frameworks. It explores using advanced audio features Mel-Spectrogram and MFCC along with LightGBM classifiers and incorporates SHAP (Shapley Additive Explanations) for model transparency and cosine similarity matrices for mimicry detection. The study identifies a dual-layered framework that attains high accuracy and robustness across datasets while ensuring adaptability to emerging attack strategies. It addresses the critical need for scalable, interpretable, and adaptive voice authentication solutions and provides a foundation for effectively mitigating evolving audio spoofing threats.

Read the paper · More papers on PaperTik