Analyzing Gender Detection in Speech in Adult Filipino Citizens
Matthew Gerard J. Abella, John Matthew L. Borja, Anne Celine F. Zuñiga, Joel C. De Goma · 2022
Being one of the most distinct and essential social characteristics and features of humans, speech is commonly used for communication. Whether it would be for the purpose of the communication of ideas, exchanging information, etc. With this, new approaches are made in order to create a more efficient and easier lifestyle for everybody. This would include gender detection or prediction for improving user experience. Affective computing, as well as artificial intelligence, is used for tasks related to recognition, including gender recognition. The study concentrates on improving gender recognition from speech, specifically Filipino adult voices. With this in mind, selecting the features that are most important in finding a speaker’s gender can be done to improve the accuracy of other algorithms as well as K-Nearest Neighbor, Random Forest, and Artificial Neural Network algorithms being used. The study then aims to explore the use of Kernel PCA instead of the classical PCA when it comes to reducing data as well as when MFCC is combined with several short term and long-term features to detect gender in speech. 250 Filipino voices from Metro Manila were gathered, cleaned, and 25 features were then extracted. Features such as MFCC, pitch, zero-crossing rate, energy, spectral centroid, spectral flux, and age were extracted. The classifiers used are Random Forest, K-Nearest Neighbor, and Artificial Neural Network. 4 models in total were created wherein Kernel PCA and PCA were applied. Out of the 4 models, those without Kernel PCA performed had a higher accuracy compared to the models that used Kernel PCA due to the information loss when lowering the dimensions. Moreover, the best performing model had an Artificial Neural Network classifier being the most accurate. This could be due to its computational power gathered from its interconnected group of nodes which is needed for speech recognition.