Age and Gender Estimation Through Speech: A Comparison of Various Techniques
Maliha Shabbir, Amjad Hussain, Maqsood Muhammad Khan · 2023
Age and gender recognition from audio signals have gained significant attention due to their potential applications in various domains, including personalized user experiences, targeted marketing, biometric identification, demographic analysis, and healthcare. This research paper presents a comparative analysis of age and gender recognition models using three different audio features: pitch, Mel-frequency cepstral coefficients (MFCC), and spectral dynamic features (SDC). For age recognition, the performance of Support Vector Machines (SVM), k-Nearest Neighbors (kNN), Random Forest (RF), Logistic Regression (LR), Gradient Boosting (GB), and Gaussian Naive Bayes (GNB) classifiers was evaluated. The SVM classifier with a linear kernel achieved the highest accuracy (96%) and precision (99%) using the SDC feature. In the case of gender recognition, SVM with the RBF kernel outperformed other classifiers, achieving an accuracy of 97.5% and an F1 score of 97.3% using the pitch feature.