Comparative Analysis of Different Neural Network Models for Speaker Gender Recognition by Voice

Ayush Porwal, Praveen Kumar Tyagi, Dheeraj Kumar Agarwal · 2023

This study examines the exploration of gender identification using voice data investigating how different deep learning models perform in this field. The models analyzed include the Artificial Neural Network (ANN) 2D Convolutional Neural Network (2D-CNN) Feedforward Neural Network (FNN) and Long Short-Term Memory (LSTM). The voice dataset used in the research consists of 20 feature columns and a label column. It undergoes steps, such, as label encoding and feature scaling. The models are carefully evaluated using important performance metrics like loss and accuracy percentages. Notably the findings highlight the strengths of each model; the ANN and FNN models demonstrate accuracy rates while the 2D-CNN model excels at capturing spatial relationships and the LSTM model focuses on temporal dependencies. The main emphasis of this work is a thorough review of the literature on speech detection and voice analysis using deep learning techniques as well as standard approaches. Beyond academia this research has implications for voice-controlled systems and security measures. The strong performance of ANN and FNN models shows promise for real world integration with applications in voice-based authentication and virtual assistants. These findings provide insights for researchers and professionals interested in developing gender recognition systems by emphasizing the importance of considering model architectures along, with their unique capabilities when working with voice data. Overall, this research adds to our understanding of the strengths and limitations of learning models when it comes to analyzing voices, for gender recognition. It provides insights for research and practical applications, in this field.

Read the paper · More papers on PaperTik