Speaker Identification Using MFCC Feature Extraction : A Comparative Study Using GMM, CNN, RNN, KNN and Random Forest Classifier

Dhanvanth Reddy Yerramreddy, Jayasurya Marasani, Ponnuru Sathwik Venkata Gowtham, Samudrala Yashwanth, S. S. Poorna, K. Anuraj · 2023

Speech recognition plays a major role in technology that enables computers and devices to understand and interpret human speech, while speaker identification is a specialized application within speech recognition that focuses on recognizing and distinguishing individual speakers based on their unique vocal characteristics. By utilizing techniques such as feature extraction, pattern recognition and machine learning algorithms speaker identification systems can match an input speech sample to a known set of speaker profiles and determine the most likely speaker’s identity. This process involves comparing the extracted features of the input speech with the features stored in a database or model. In this project we are performing a comparative study which mainly focuses on extracting features through MFCC feature extraction for all models in common and then the implementation of identifying speakers using Gaussian Mixture Model, CNN, LSTM in case of RNN, k-Nearest Neighbour and a Random Forest Classifier.A comparative study on various types of models helps us to pave way on analyzing these to implement one of the best among these in day to day life. Among all the models, GMM performed very well with an accuracy of 98.68% followed by LSTM with an accuracy of 95.77%.

Read the paper · More papers on PaperTik