Multilingual and gender classification of speech using acoustic and deep learning methods
Kunal Ghosh, Aviruk Basak, Sagnik Bhattacherjee, Partha Pratim Mohanta, Saiyed Umer · Procedia Computer Science · 2025
This study investigates the classification of both language and gender using speech signals from five distinct languages. A compre- hensive set of acoustic features is extracted, including Mel-Frequency Cepstral Coefficients (MFCC), Zero Crossing Rate (ZCR), Root Mean Square Energy (RMSE), statistical descriptors, and Chroma Short-Time Fourier Transform (Chroma STFT). Four ma- chine learning classifiers are evaluated on a merged dataset comprising SLR 41, 42, 43, 44, and SpeechOcean762. Experimental results reveal that features obtained from a 1D Convolutional Neural Network (CNN), when used with a Random Forest classifier, yield an F1-score of 98.49% for gender classification. For language identification, a Support Vector Machine (SVM) achieves an F1-score of 98.75%. In contrast, non-fine-tuned Wav2Vec embeddings produce lower performance, with F1-scores between 74% and 77%, while handcrafted acoustic features reach F1-scores of 96% for both tasks. These findings emphasize the continued effec- tiveness of traditional signal processing techniques in multilingual speech classification and highlight the necessity of fine-tuning when utilizing pre-trained deep learning models. The application of Principal Component Analysis (PCA) for feature reduction results in a slight decline in classification performance across all tested models.