Gender Recognition and Classification of Speech Signal

Icsmdi Submitter, Bhagyalaxmi Jena, Anita Mohanty, Subrat Kumar Mohanty · SSRN Electronic Journal · 2021

Human speech comprises different sounds for the purpose of speaking, singing, and expressing emotions and ideas, etc. which is generated by the vibration of the vocal cord of human beings. The rate of vibration of vocal cord is termed as pitch which is a frequency domain parameter. Vocal tracts of males are mostly longer in comparison to vocal tracts of females which leads to differences in human speech for males and females. With increase in demand of Human Computer Interaction (HCI) systems, speech processing plays an important role in enhancing HCI systems. The development of gender recognition systems finds its applications in gender based virtual assistants, telephonic surveys and voice controlled automation systems. Research works were mostly done in either frequency domain or time domain. In this work, speech signal analysis was based on both time and frequency domain. Different speech parameters were generated by short-time, statistical and spectral analysis. The differences in parameters was used as a working principle for the gender model to recognize the gender of the unknown user. The classifier model based on Genetic Algorithm, Gaussian Mixture Model (GMM) is known to have an accuracy of 70% with complex training. So, the classifiers used in this work were KNN and SVM. After training and testing, the accuracy of the system was found to be 80 percent on average.

Read the paper · More papers on PaperTik