A Hybrid Model of CNN-SVM for Speakers’ Gender and Accent Recognition using English Keywords

Yeshanew Ale Wubet, Kuang‐Yow Lian · 2021

Nowadays, the speakers’ accent recognition, speech to text conversion, and their applications are becoming popular research areas all over the world. This paper proposes a hybrid model composed of Convolutional Neural Network (CNN) and Support Vector Machine (SVM) for gender, accent, and keyword classification. The result of the hybrid model is better than just using CNN or SVM. It is well known that the training of the hybrid model will be more complicated than the training of pure CNN or SVM. The CNN extracts features from a spectrogram image representation of speech and SVM is applied to extracted features as a classifier. The fusion model of CNN-SVM converges fast and reduces the overfitting problem unlike to CNN model alone. The result shows that the proposed system carried out multiple tasks at the same time and achieved high recognition accuracy.

Read the paper · More papers on PaperTik