Speaker identification using auditory modeling and vector quantization
Konstantina Iliadi, Stefan Bleeck · The Journal of the Acoustical Society of America · 2015
Speaker identification (SID) aims to identify the underlying speaker(s) given a speech utterance. SID systems can perform well under matched training and test conditions but their performance degrades significantly because of the mismatch caused by background noise in real-world environments. Achieving robustness to the SID systems depends very much on the front-end (or feature extractor), which is the first component in an automatic speaker recognition system. Feature extraction transforms the speech signal into a compact representation that is more discriminative than the original signal. We present on our poster a new system where the parametrization of the speech is based on an auditory model called Auditory Image Model (AIM). Two experiments were performed for two different sets of speakers. Experiment 1 identified the most informative regions of the auditory image that can indicate speaker recognition. Experiment 2 consisted of training 10 and 60 speakers using clean speech and testing those two groups using speech in the presence of babble noise of eight speakers for 5 SNRs. The results suggest that the extracted auditory feature vectors led to much better performance, i.e., higher SID accuracy, compared to the MFCC-based recognition system especially for low SNRs.