Enhancing Voice Pathology Detection with Ensemble Stacking and Machine Learning

M. L. Rakhil, Aman Sirohi, R. Sriviswa, G. Jyothish Lal · 2024

Classification of the voice pathology plays an important role in both medical field as well as speech processing because voice affects communication, interpersonal relations, and mental well-being. This research work aims at adopting a new technique of Machine learning (ML) that will help in the diagnosis of more serious diseases such as hyperfunctional dysphonia, laryngitis, as well as vox senilis for better patient care and management and ultimately help to cut down the costs incurred in patient health care management. Additive White Gaussian Noise (AWGN) and time stretching was applied to introduce variations to the dataset, while the OpenSMILE toolkit was used to extract key acoustic features from the data. Hyperparameters were optimized with the help of the Optuna library. The best performance of the classifiers was obtained using ensemble stacking, where the final Random Forest classifier achieved an accuracy of 91%. This work also revealed that the meta-model is superior to both the Deep neural networks (DNN) and 1D Convolutional neural network (CNN) in terms of accuracy and precision. It also emphasises on the possibility of recognising and treating voice disorders at earlier stages, thus improving quality of life of patient treated with this disorder.

Read the paper · More papers on PaperTik