Pathological Voice Detection Using Transfer Learning Methods
Zhang Yihua, Xincheng Zhu, Yuanbo Wu, Zhang Xiaojun, XU Yi-shen, Zhi Tao · 2021
Pathological voice detection has obtained great progress on recognition rate. However, these results are all achieved in single database. In cross-database detection, the accuracy of detection tends to drop. Transfer learning, which has been proven to be able to address cross-domain recognition, is applied to pathological voice detection. The aim of transfer learning is to find a mapping matrix, which maps the source and target data into a common feature subspace to reduce the difference of the feature distribution between different databases. Meanwhile inner-class and inter-class distance are applied to transfer learning to keep the features of the original space as much as possible after dimensionality reduction. Three dimensional scatter plots are used to evaluate the our proposed methods. Finally, after transfer learning methods, the accuracy is increased by up to 9.14%. Experimental results demonstrate that the accuracy of cross-database pathological voice detection can signification improve by transfer learning methods.