Enhancing Gender Detection in Speech using Generative Adversarial Networks and Autoencoder Architecture
Mandar Pramod Diwakar · Panamerican mathematical journal. · 2025
Gender detection in speech analysis plays an important role in voice-based applications, human-computer interaction, and biometric systems. The present study used Generative Adversarial Networks (GANs) to use Mel-Frequency Cepstral Coefficients as input features in gender detection. Using a database of labeled audio recordings, the authors extracted 13 MFCCs per sample to effectively capture the spectral and temporal characteristics of speech. These features were averaged across time frames and standardized to ensure consistency during the training and testing phases. The resulting 13-dimensional feature vectors were processed using a GAN architecture, with the discriminator trained to classify gender based on the input features. Experimental validation on a dataset of audio samples demonstrated the system's efficiency, achieving an average accuracy of 94.5% in gender classification tasks. The model processes a 10-second audio sample (approximately 160,000 data points) in 0.02 seconds on a standard CPU, showing the computational efficiency and scalability of the model. A classification threshold of 0.5 was used, where predictions above 0.5 indicated female speakers and below 0.5 indicated male speakers. This pipeline does highlight the prospects of using GANs and MFCC features to present effective and reliable gender detection systems. The approach provides a scalable, reproducible framework that could easily be used for deployment in different applications in the real world, like virtual assistants or speech-based analytics.