Optimized artificial neural network for vocal gender expression classification
Mohamed Mansoor Roomi Sindha, K. Mohana Priya, Uma Maheswari Pandyan, M. Sreenivasan · ETRI Journal · 2025
Abstract Gender expression is a complex and multidimensional characteristic that goes beyond traditional binary models, especially regarding vocal attributes. Existing voice gender recognition systems often rely on binary classification, which fails to capture the subtle spectrum of gender expressions in natural speech. To address this limitation, we introduce the FEMASC dataset, which is a novel resource that categorizes voice signals into five classes (absolute male, strong male, ambiguous, strong female, and absolute female) reflecting varying degrees of masculinity and femininity. In addition, we propose an optimized artificial neural network that uses feature fusion and Bayesian hyperparameter tuning to effectively model subtle variations in vocal features. Our approach not only overcomes the constraints of binary classification but also achieves robust performance with a precision of 86.8%, recall of 86.1%, F 1 score of 86.2%, and accuracy of 82% on the FEMASC dataset. These results demonstrate that a multiclass framework supported by a carefully designed dataset and advanced modeling techniques offers significant improvement in voice gender expression recognition.