A Robust and Lightweight CNN-Transformer Model for Audio Deepfake Detection in Indian Languages
Manish Gaikawad, Soma Niloy Ghosh · 2025
The proliferation of audio deepfake technology poses significant threats to digital security, particularly in multilingual contexts such as India, where diverse languages are widely used. This paper presents a robust and lightweight CNN-Transformer hybrid model designed for audio deepfake detection in major Indian languages, including Marathi, Hindi, Tamil, Telugu, Malayalam, and Kannada. The proposed model addresses two critical challenges in deepfake detection: domain generalization and computational efficiency. By integrating multiple feature types—Mel spectrograms, LFCC, and Wav2Vec—and leveraging Convolutional Neural Networks (CNNs) for spatial feature extraction, Transformer layers with multi-scale attention for temporal dependencies, and a spectral-temporal gating mechanism for feature fusion, the model achieves superior performance across diverse audio environments, including low-quality and noisy data. Additionally, the model is optimized for edge devices using pretraining, quantization, knowledge distillation, and pruning, ensuring real-time performance with minimal accuracy loss. Evaluated on a novel dataset of nearly 10,000 samples per language, alongside ASVspoof 2021 and in-the-wild datasets, the model achieves a test accuracy of 98.07% (validation accuracy of 98.62%) with precision and recall scores of 0.98 and 0.98 on the test set, respectively. The study also analyzes trade-offs between model size, inference speed, and detection accuracy, offering insights into practical deployment. This work advances audio forensics by providing a scalable, efficient, and language-agnostic solution to counter audio deepfakes in resource-constrained settings.