Speaker Gender Recognition Based on Semi-Supervised Learning
Zheyan Zhang, Richard Z. Li, Kewei Chen · 2024
Determining a speaker's gender from speech signals is essential for intelligent human-machine interaction. Significant progress has been made in supervised learning for speaker gender recognition. However, effective performance relies heavily on large labeled datasets. To address this, a semi-supervised learning framework is introduced to fully leverage vast amounts of unlabeled speech data for gender recognition. First, contrastive learning is employed to enhance and augment a relatively small set of labeled audio data, training an encoder capable of distinguishing speech genders. Next, a fuzzy clustering algorithm is utilized to generate gender pseudo-labels from a large pool of unlabeled data. Finally, the speaker gender recognition model is iteratively trained by using these pseudo-labels. Experiments and analyses on the Common Voice English dataset demonstrate that the model attains a mean accuracy of 97.4% in gender recognition tasks.