Modifying deep neural network structure for improved learning rate in speakers' age and gender classification

Zakariya Qawaqneh, Arafat Abu Mallouh, Buket D. Barkana · 2016

Speaker's age and gender classification is one of the most challenging problems in the field of acoustic recognition. Although many studies have been done to obtain better results, the classification accuracies are still not satisfactory. Motivated by the success in the deep learning techniques in speech processing field, we developed a DNN architecture to classify speakers' age and gender. In this work, we propose a new data model and we make modifications in GB-RBM and the BB-RBM architectures. This work shows that DNN can be trained by using age and gender real valued data and it will have a solid expressive and distinctive capability to classify the speaker's age and gender if the DNN is skillfully initialized and trained. Experimental results showed that the proposed DNN model and modifications achieved faster learning. To evaluate the proposed model, several experiments were conducted. The dimensional grey scale displays for the activation probability of hidden units, weights and biases histograms, and the MSE metrics were used.

Read the paper · More papers on PaperTik