Subthreshold Depression Detection With Text-Guided Multimodal Learning

Yanrong Guo, Youwei Guo, Bingxin Yang, Jingjing Wu, Shijie Hao, Richang Hong · IEEE Transactions on Computational Social Systems · 2025

Depression, a widespread global mental health problem, affects millions of people annually, making early detection of subclinical depression crucial for timely intervention. Current automatic depression detection (ADD) methods, valuable for diagnosis, often neglect subthreshold populations and face difficulties in extracting diagnostic data from long sequences of multimodal information. These methods also inadequately leverage text modality, which is less noisy and information-rich compared with other modalities. Furthermore, existing datasets for depression research are often too small, limiting the generalizability of developed methods. To address these issues, this article proposes a new approach for detecting depression in subthreshold populations. For long-sequence samples in the field of depression, we construct an autoencoder that compresses along both temporal and feature dimensions, aiming to extract the most compact and effective features from the samples. To exploit the text modality’s advantages, we integrate the RoBERTa pretrained model with an attention mechanism for high-quality text encoding. We then develop a text-guided multimodal fusion (TGMF) module, using text encoding as an anchor for guiding audio and video modality encoding, ensuring multimodal alignment. Additionally, contrastive learning is applied to discern differences between classes, enhancing the model’s generalizability. Our method demonstrates superior performance in the tasks of detecting depression and identifying subthreshold populations on the E-DAIC and MMDA datasets.

Read the paper · More papers on PaperTik