Pathological Voice Detection From Sustained Vowels: Handcrafted vs. Self-supervised Learning
Bagus Tris Atmaja, Akira Sasou · 2025
Pathological voice detection aims to detect voice disorders from speech samples. With the recent development of self-supervised learning (SSL), most studies in the past years of voice disorder employ that technique. Handcrafted features may suggest better prediction since they contain physical information and are more interpretable than SSL. We evaluated different handcrafted acoustic features and SSL approaches for pathological voice detection tasks using sustained vowels. We extracted 88 and 39 dimensional handcrafted features using openSMILE and Praat feature extractors. For SSL, we evaluated wav2vec 2.0, HuBERT, and WavLM, both for feature extractors and finetuning. Results showed that handcrafted features are consistently competitive with SSL features. An ensemble model combining handcrafted and SSL features achieved the best performance with an F1-score of 0.8739 on the test set, outperforming previous studies on the same dataset under the test set. This finding suggests that handcrafted features are still competitive for voice disorder detection tasks, and combining them with SSL features can further improve performance.