Speech recognition of speaker identity based on convolutional neural networks
Hangdong An · 2023
Recent advances in the field of speaker recognition (SR) have shown very accurate and high-performance algorithms. When the amount of data is small, its performance drops dramatically [1]. Today, when testing and training involve only small amounts of speech data, identifying the speaker remains a key consideration, as real-life applications can only access speech data for a limited period. Speech recognition-based security systems are one of the main areas of research. In this paper, we will use a graphical CNN algorithm for speaker recognition and compare the effect of different times of speech on the accuracy of the model. The results show that better results can be obtained by using the speech feature data with a deep learning CNN model, with an average accuracy of 86.3%. In addition, the accuracy of speech recognition can be better improved by comparing multiple sets of speech segments with smaller results in this study.