The Effect of Balanced and Imbalanced Dataset on Voice Pathology Detection Using Online Sequential Extreme Learning Machine (OSELM)
Nurul Fariesya Suhaila Md Sazihan, N. M. Abdul Latiff, Nik Noordini Nik Abd Malik, Syamsiah Mashohor, Khaled Abdulaziz Alaghbari, Fahad Taha AL‐Dhief · 2024
Voice pathology detection is crucial for early diagnosis and treatment of vocal disorders. This research studies the effect of balanced and imbalanced datasets on the accuracy of voice pathology detection using Online Sequential Extreme Learning Machine (OSELM) on the Saarbruecken Voice Database (SVD). The dataset's balance has been identified as having a significant influence on the performance of the machine learning models. In this work, the voice signals from vowel /a/ of pathology and non-pathology classes are taken from SVD. The dataset is divided into two category which are balanced and imbalanced dataset. Then, the features of the voice signals are then extracted using Mel-Frequency Cepstral Coefficient (MFCC) and fed into classifier called OSELM. The result of balanced and imbalanced dataset is assessed for its accuracy. The results showed that the model achieved an accuracy of 66.25% with the balanced dataset compared to 48% for the imbalanced dataset. These findings highlight the importance of addressing dataset balance to enhance the reliability and effectiveness of voice pathology detection systems.