Multi-Feature Fusion-Based Speech Disorder Classification Using MobileNetV3-EfficientNetB7, Linformer-Performer, and SHAP-Aware XGBoost

Abdul Rahaman Wahab Sait, Suresh Sankaranarayanan, P. Gouthaman · IEEE Access · 2025

Traditional speech disorders (SD) detection relies on subjective analysis, resulting in inconsistent outcome. Direct voice classification lacks effective approaches to capture temporal dependencies. Machine learning (ML) models face challenges in extracting the complex temporal and spectral variations in speech signals. Advanced deep learning (DL) and transfer learning techniques offer a foundation for early screening of SD. However, the lack of interpretability reduces the generalization capabilities of these models. Addressing these shortcomings is essential in order to improve the accuracy of SD detection and clinical trustworthiness. Thus, we introduce a novel image-based SD classification model to classify healthy and pathological speech with high accuracy and robustness. We transform raw speech signals into Mel-Spectrograms to overcome the limitations of direct voice classification. To facilitate the model’s interpretability, the statistical and handcrafted acoustic features are extracted from the raw speech signals. Hybrid MobileNet V3-EfficientNet B7and Linformer-Performer are employed to extract diverse features from the Mel-Spectrograms. An attention-based feature fusion is used to identify critical features indicating the SD patterns from the extracted features. We fine-tune XGBoost classifier to classify the healthy and pathological speech. SHapley Additive exPlanations (SHAP) values is employed to offer valuable insights into the model’s decisions. The proposed model obtains an exceptional performance on two benchmark datasets. On the Saarbruecken Voice Database (SVD), it achieves an accuracy of 98.9% with loss of 0.09. It yields a remarkable generalization accuracy of 98.2% on the VOICE dataset, outperforming the state-of-the-art models. In addition, it contributes a significant advancement in SD detection, setting the stage for future research endeavors.

Read the paper · More papers on PaperTik