Prediction of Age from Speech Features Using a Multi-Layer Perceptron Model

Sriram Ravishankar, Prasanna Kumar M K, Vinay Vasanth Patage, Sourabh Tiwari, Saksham Goyal · 2020

Age estimation from speech has recently received enormous interest for its plethora of applications such as personal voice assistants, user-profiling, targeted marketing, call-routing and criminal tracking. There is critical demand for a system that can estimate the age of the speaker quickly and accurately. In this paper we discuss the prediction of age group from speech features using a Multi-Layer Perceptron (MLP) architecture. Features like Mel Frequency Cepstral Coefficients (MFCC), formants, pitch, chroma values etc. are extracted from every speech sample and are then selected using Principal Component Analysis (PCA) and Redundant Feature Elimination (RFE) techniques. The training and testing are done on a repository from the Mozilla Speech database called Common Voice, consisting of around seventy-five thousand speech samples from different age groups. The implemented model gave us a training accuracy of 87.59% and testing accuracy of 85.68%. We have explained the approaches used in the existing solutions and their drawbacks. Having compared the performance from state-of-the-art solutions on this dataset, we have shown that our proposed model provides enhanced age group classification accuracy.

Read the paper · More papers on PaperTik