Automatic gender detection

Binodini Tripathy, Niva Das, P R Nayak, Unnayan Raj, S.K. Datta · 2025

Automatic gender detection problem has gained significant attention in recent years due to its numerous diverse applications such as human-computer interaction, speech/speaker recognition, biometric security, biometric attendance system, etc. The process of detection, when carried out using voice samples, eases the process of implementing security protocols with better accuracy than using facial recognition approach. Gender detection can be viewed as a classification problem in which the system classifies the input audio signal into two categories: i.e., male and female. Automatic speech recognition systems with built-in gender-specific models produce higher recognition rates as compared to gender-independent recognition systems. Since speech samples contain unique linguistic and paralinguistic features such as gender, age, emotional state, etc., extraction of these features that can be used in Machine learning (ML) and Deep learning (DL) models for developing gender detection subsystem seems promising and effective. Though different time domain and frequency domain features have been used by researchers, in this work, we have explored only two frequency domain features namely, pitch and Mel-frequency-cepstral-coefficients (MFCC) for identification of gender. We have considered a popular emotion dataset namely RAVDESS for this purpose and used fivefold cross-validation technique for training and testing of these datasets. Five ML classifiers, i.e., Support vector machine (SVM), Logistic regression, Decision Tree (DT) and Naïve Baye’s have been trained and tested after preprocessing the voice samples followed by feature extraction. We have also employed Convolutional Neural Net (CNN) for the same task with MFCC features as input on Common Voice Dataset available from Kaggle. The performance of all these classifiers is presented in terms of accuracy, precision, and recall. Among ML classifiers, SVM and Logistic Regression produce 99% accuracy on RAVDESS dataset while CNN also produces 99% accuracy when fed with MFCCs as input. On Mozilla Common Voice dataset with MFCC input, SVM yields 95.4% accuracy, whereas CNN yields 96.6% accuracy. The experimental results are not only encouraging but also vital for real-time implementation.

Read the paper · More papers on PaperTik