Prosody Disentanglement with Self-Supervised Speech Representation for Detecting Depression
Bubai Maji, Rajlakshmi Guha, Aurobinda Routray, Shazia Nasreen, Debabrata Majumdar · 2025
Speech signals offer valuable insights into mental health, especially in diagnosing depressive disorders. Human speech encompasses various components, such as semantic content, speaker identity, and prosodic information. A key challenge remains in disentangling prosodic information from other components due to their intrinsic association and the need for robust depressive disorders detection systems. This paper aims to address the disentanglement of prosody information by leveraging self-supervised learning (SSL) model for depression detection. Specifically, our model, DepAug, captures both verbal and non-verbal characteristics through a semantic encoder and a prosody encoder. The decoder component then enables the generation of synthetic samples to address data imbalance issues. Experimental results on our Bengali dataset show that the model accurately captures general prosodic characteristics that can adapt to diverse emotional speech contexts. Additionally, we train a self-supervised speech model using DepAug’s data augmentation, showing that it outperforms state-of-the-art supervised and self-supervised approaches.