Utilizing Artificial Neural Networks and Mel-Frequency Cepstral Coefficients for Gender Identification from Voice Data

Beeram Bhanu Teja, Madan Lal Saini, Edupalli Greeshmanth Kumar, Syed Abbas Ali · 2024

Gender recognition from voice has drawn a lot of interest because of its wide-ranging many uses in areas including speech processing computer-human exchange, and voice-based personalization. Gender recognition through voice analysis is a critical aspect of speech processing, finding applications in security systems, voice assistants, and user profiling. This study explores the implementation of Artificial Neural Networks (ANN) for precise voice-based determination of gender features. The characteristics that were taken out include Mel Frequency Cepstral Coefficients (MFCCs), pitch, formants, and other relevant acoustic parameters. The dataset utilized is sourced from the common Voice library that the Mozilla Speech Database gives accessibility. The research employs the Common Voice dataset from the Mozilla Speech Database. The MFCC helps to convert the voice data into 26 numerical features. The accumulation of data comprises of 193,933 voices in the training set and 50,000 voices within the testing group. The analysis involves training an ANN model with 100 epochs on MFCCextracted features from real voices. The achieved training accuracy stands at 87.0%, and the testing accuracy of 80.95% indicates the efficacy of the model in discerning gender from voice samples. This research contributes to the advancement of voice-based gender recognition technology across various applications. However, further exploration is needed to enhance the model's generalization capabilities and address potential limitations, particularly in accommodating diverse speech patterns and environmental variations.

Read the paper · More papers on PaperTik