Pathological Voice Recognition Based on Multi-channel Feature Fusion Network

Yi‐Ming Fan, Penghui Zhao, Yihui Zhang, Haibin Yuan, Jianxun Lv · 2025

Pathological voice recognition technology based on voice signal analysis for voice quality assessment is an essential clinical guideline for more and more people suffering from voice disorders. However, most existing algorithms only consider the binary classification of healthy and morbid voices, and the temporal characteristics of the voice signal are ignored. In this paper, a multi-channel feature fusion network (MFFN) based on a voice recognition algorithm is proposed for detecting healthy voice, neurological voice, and physiological voice by the semantic and temporal features. The algorithm’s performance is experimentally validated using the Saarbrucken Voice Database (SVD), and the model achieves an accuracy of over 91%. The experimental comparison between fused features and single features is also presented, showing that fusion features are significantly better than single features in detecting pathological voices.

Read the paper · More papers on PaperTik