SPEECH BIOMETRICS: A Comprehensive Deep Learning-based Speaker Identification System
Pooja Shetty, Ryan Rodricks, Sahil Malgundkar, Hirthik Pamnani, Shubham Katke · 2023
Speaker recognition is a fascinating research area that has grown significantly in the last few years. The recognition rate of this technology depends on several factors such as pitch, frequency, handset type, and background noise, among others. Automatic Speaker Recognition (ASR) is the process of identifying and verifying a person based on their speech waves and speaker-specific information. It is used in various fields, including access control for different voice-based services, such as phone dialing, voice mail, database services, telephone shopping and banking services by telephone. Speaker recognition is also used in the security monitoring of remote access to computers and confidential information. Speaker recognition technology plays an important role in forensics. It is a valuable tool that is used to identify suspects or witnesses based on their speech recordings. Therefore, understanding the basics of ASR, including feature extraction techniques and speaker modeling, is crucial to achieve accurate speaker recognition. The objective of the paper is to evaluate and compare the performance of Siamese Neural Networks (SNN) and Gaussian Mixture Models (GMM) in the context of speaker recognition. Specifically, the paper aims to assess their effectiveness in distinguishing between four different speakers using audio inputs and preprocessing techniques, with a focus on achieving low Equal Error Rates (EER). The paper employed a comparative methodology where it utilized two different audio inputs from four distinct speakers. Preprocessing was performed on the audio data using the MFCC (Mel-Frequency Cepstral Coefficients) algorithm, followed by the removal of silent segments. The processed audio data were then input into both a Siamese Neural Network and a Gaussian Mixture Model. The Siamese Neural Network and GMM were used as the models for speaker recognition. The SNN achieved an EER of 4.88%, while the GMM achieved a slightly higher EER of 6.55%. In summary, speaker recognition has several important applications in our daily lives, and the accuracy of the technology depends on various factors. ASR is a fascinating field of research that continues to evolve, and understanding its basics is essential to achieve accurate and reliable results.