Feature extraction of audio data for speaker’s gender classification

Assila Yousuf, David Solomon George · Journal of Physics Conference Series · 2025

Abstract Embracing deep learning techniques in speech processing has revolutionized the field, offering new possibilities for improving speaker classification and gender identification. This research examines the application of sophisticated neural network frameworks, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs), for the extraction of characteristics from audio signals. By employing these deep learning models, we aim to automate the extraction of intricate features using raw audio data, which in turn boosts accuracy of the classification. We investigated the influence of different audio features, including spectrograms and Mel-frequency cepstral coefficients (MFCCs), regarding the effectiveness of speaker and gender classifiers. The outcomes of the experiment reveal that deep learning-based feature extraction significantly outperforms traditional methods, offering robust and precise classification results.

Read the paper · More papers on PaperTik