Robust Automatic Speaker Identification System Using Shuffled MFCC Features
Mahdi Barhoush, Ahmed Hallawa, Anke Schmeink · 2021
Speaker Identification using deep learning is an ongoing active topic that has wide applications in voice authenticated and activated systems, forensic investigations, security, and remote transactions. Mel frequency cepstral coefficients (MFCC) feature-based speaker identification systems are proven to per- form very well with clean speech conditions. However, there is rapid degradation in the robustness with adverse conditions and short speech duration. To overcome this limitation, we propose a new speaker identification pipeline system grounded on novel MFCC based features called shuffled MFCC (SHMFCC) along with a new data augmentation approach. Moreover, the system uses a simple and efficient deep neural network model with variable-duration input speech signals and is proved to perform remarkably at different datasets with different environmental and noise conditions and for a high number of speakers.