Voice Gender Classification using Deep Learning and MFCC Features
Ankit Kumar Singh · 2025
This paper presents a comprehensive approach for classifying human voice gender using deep learning techniques with Mel-frequency cepstral coefficients (MFCCs) as features. We utilized a balanced subset of Mozilla's Common Voice dataset and implemented both feedforward and convolutional neural networks for gender classification. Extensive preprocessing, MFCC matrix generation, CNN model design, and frontend integration via Streamlit are covered. The model achieved an accuracy of 84%, with robust precision and recall metrics. This system is further enhanced with a user-friendly web interface supporting file upload and real-time audio recording for live predictions.