Emotional and cognitive assessment from speech: an investigation of machine learning solutions towards trustworthy clinical applications

Edward Lázaro Campbell Hernández · 2025

The escalating global prevalence of mental health disorders, including Major Depressive Disorder (MDD) and Alzheimer’s Disease (AD), underscores the urgent need for accessible, non-invasive diagnostic tools. Speech, as a biomarker integrating cognitive, emotional, and physiological information, offers a promising avenue for early detection. This thesis investigates machine learning (ML) solutions to automate the assessment of emotional and cognitive states through speech analysis, focusing on developing trustworthy clinical applications. Using a diverse experimental framework comprising five datasets (ADReSS, AcceXible-MCI, DAIC-WOZ, RADAR-MDD, and Androids), this work evaluates paralinguistic (e.g., acoustic prosody, Low-Level-Descriptors, self-supervised features) and linguistic (e.g., GloVe, lexical diversity) features. Key findings reveal that Transformer-based models, leveraging semantic content, achieve state-of-the-art performance in detecting MDD and cognitive decline, while self-supervised acoustic features demonstrate cross-lingual robustness, particularly in low-resource settings. Personalized, speaker-dependent frameworks significantly outperform generalized models, highlighting the importance of tailoring systems to individual variability in symptom expression. The study identifies noun/adverb ratios and emotional valence as statistically significant indicators of cognitive impairment and MDD, respectively. Furthermore, ensemble strategies combining top-performing models enhance detection accuracy (e.g., 90.04 % AUC for MDD detection in DAIC-WOZ). A critical contribution is the integration of explainable AI (XAI) techniques, such as SHAP values, to elucidate model decisions, fostering clinician trust by aligning feature importance with clinical markers such as lexical impoverishment. The thesis underscores the viability of speech-based ML tools as scalable screening aids in healthcare, addressing ethical considerations through transparent methodologies. Future work will explore longitudinal monitoring and multimodal integration to refine diagnostic precision and adaptability across languages and demographics

Read the paper · More papers on PaperTik