Audio Spoof Detection Using Inverse Gammatone Cepstral Coefficients and Multi- Layer Perceptron Classifier
Abhishek Nandal, Mohit Dua · 2024
The rise of synthetic audio and voice cloning technologies has necessitated the development of robust methods for detecting audio spoofing. Accurate and reliable spoof detection is critical for maintaining the integrity of voice-based authentication systems and preventing malicious misuse of synthetic speech. This paper presents a approach to audio spoof detection utilizing Inverse Gammatone Cepstral Coefficients (IGTCC) for feature extraction and a neural network for classification. The IGTCC method captures essential audio characteristics, enhancing the model's ability to differentiate between genuine and spoofed audio. Our system is evaluated using the Fake or Real (FoR) dataset, which comprises a balanced collection of both authentic and spoofed audio samples. The dataset includes various spoofing techniques, providing a comprehensive benchmark for testing our model's effectiveness. The proposed neural network architecture is trained and validated on this dataset, demonstrating its efficacy in accurately identifying spoofed audio samples. The system achieves a remarkable Equal Error Rate (EER) of 1.6%, underscoring its potential for deployment in real-world applications. This significant improvement over existing methods highlights the effectiveness of our approach in addressing the challenges of audio spoof detection.