Proposal and Implementation of Neural Network-Based Approach for Secure Speaker Recognition Using Spectrograms
Arindam Bindlish, Vaibhav Gujral, Rosy Madaan, Praveen Kumar · 2025
In this era of Artificial Intelligence, research on Speaker Recognition systems has taken the forefront. In this paper, an efficient deep learning pipeline has been proposed. A total of 7500 audio samples from Kaggle’s speaker recognition dataset have been used to generate Spectrograms, Mel Spectrograms, and Mel-Frequency Cepstrum Coefficients (MFCC), which in turn have been used to train a ResNet 18 Deep Neural Network from scratch. Using different regularization techniques such as dropout, early stopping, and weight decay, high accuracies were obtained. The results generated from the ResNet by using Mel spectrograms, classic spectrograms, and MFCCs were compared alongside each other to check which one among them gave the best testing accuracy. The proposed approach is designed and implemented with additional feature enhanced by providing a secure feature selection system from audio signals. To implement this, the proposed system is equipped with AES security algorithm on splitting of data set into ratio of 60:30:10. The Classic Spectrograms predicted the speakers with 92.96% accuracy. Mel Spectrograms were comparatively more efficient with 93.75% testing accuracy. The results obtained from the model trained with the MFCC features were the best of all, at 95.2%. Further, the usage of the Advanced Encryption Standard (AES) algorithm has also been proposed for a secure speaker recognition task.