Text-independent Speaker Recognition: A Deep Learning Approach
Arman Shirzad, Ali Akbar Nasiri, Razieh Darshi, Zohreh Safarpour, Razieh Abdollahipour · 2024
In this study, we introduce an artificial neural network (ANN) model specifically designed for speaker recognition and compare its performance and accuracy with that of the established Keras speaker recognition model. Our investigation utilizes two distinct datasets: the first comprises voice recordings in three different languages from both male and female speakers; the second is the widely recognized LibriSpeech dataset, a comprehensive corpus of English speech derived from 50 individuals reading audiobooks. The proposed model leverages the Fast Fourier Transform (FFT) to analyze the input data, which consists of labeled .wav files. These files undergo a preprocessing stage before being input into the ANN for training. Our results indicate that while the proposed model outperforms the Keras model on the first dataset, it exhibits marginally lower performance on the LibriSpeech dataset. This finding suggests that even well-established models like those provided by Keras may not universally deliver optimal results across different datasets.