Age, Gender and Emotion Recognition by Speech Spectrograms Using Feature Learning
Yash Karbhari, Vaibhav D. Patil, Pranav Shinde, Satish Kamble · 2023
Speech signals are complex and dynamic signals that carry important linguistic information and ideas. The origins of human speech date back to around 300,000 to 500,000 years ago, coinciding with the emergence of anatomically modern humans. Despite the importance of speech in human communication, the accurate classification of the human voice in digital media remains a challenging task in pragmatic applications. To address the previously mentioned challenge, this research paper presents a state-of-the-art technique for identifying various speaker characteristics, including age, gender, and emotion, by analyzing sampled audio. For age and gender classification, eight different algorithms were experimented with to achieve maximum accuracy, where K-nearest neighbor (KNN) performed with the highest accuracy, and for emotion recognition, a simple five-layered CNN module was developed. Relevant features were extracted from the audio samples using different techniques to identify and classify the different speaker characteristics. Multiple datasets were utilized to perform various tasks, including the Common Voice dataset for age and gender classification, and the combination of CREMA-D and RAVDESS datasets for emotion recognition. The system's accuracy is evaluated using test samples, resulting in 79.57% and 93.26% accuracy for age and gender classification respectively, and 98% accuracy for emotion recognition. Overall, the proposed technique aims to contribute to the development of a novel and valuable tool for the audio-based analysis of human characteristics. The system has potential applications in various fields, including healthcare, psychology, and security.