A Deep Multimodal Voice Pathology Classifier with Electroglottographic Signal Processing Capabilities
Ioanna Miliaresi, Aggelos Pikrakis, Kyriakos Poutos · 2022
In this paper, we present a study on the integration of multimodal data into a modular deep learning architecture, for the task of automatic classification of voice pathology, with an emphasis on the role of electroglottographic signals as a new modality. The proposed architecture fuses audio descriptors, categorical medical data and features derived from electroglottographic signals into a modular deep learning architecture. More specifically, a cascade of convolutional layers processes a sequence of short-term audio feature vectors consisting of Mel-Frequency Cepstral Coefficients, their derivatives and Mel filterbank outputs. Furthermore, medical records and mid-term audio features are processed by a standard feed-forward branch and, finally, in a parallel, convolutional branch, ”wavegrams” of the glottal chords are processed as a new modality to increase system accuracy. Our study focuses on the sustained vowel /a/ in a subset of the Saarbrucken voice database. The resulting problem is class-imbalanced and the gathered dataset includes healthy subjects and three types of pathology types, namely hyperfunctional dysphonia, laryngitis and recurrent laryngeal nerve paralysis. The experimental results indicate that the integration of various modalities and in particular the inclusion of electroglottographic signals, increases the system’s classification capabilities, yielding a state-of-the-art classification accuracy of 89.3%.