A Research on Multimodal Depression Detection Based on Semantic Enhancement and Lightweight Optimization of CubeMLP

Jinlin Li, Lifeng Yin · Recent Patents on Mechanical Engineering · 2025

Introduction: Depression is one of the most prevalent mental disorders worldwide, and its accurate and objective detection has become a critical challenge in both medical and computer science domains. Existing multimodal depression detection methods often suffer from weak crossmodal semantic alignment, arbitrary feature fusion, and limited validation of generalization performance. Methods: To address these issues, we propose a novel semantic-enhanced multimodal depression detection framework based on a BERT-BiGRU backbone, graph-based semantic fusion, and an optimized CubeMLP classifier. Our method first uses BERT and BiGRU to extract textual, acoustic, and visual features, which are then enhanced through a multi-head self-attention mechanism to improve intra-modal contextual understanding. A cross-modal multi-head attention module is introduced to enable semantic alignment across modalities, and a Graph Neural Network (GNN) is used to model inter-modal relationships via adaptive message passing. To reduce computational complexity while maintaining strong representational capacity, we integrate Depthwise Separable Convolutions (DSC) and introduce ReZero residuals and Mish activation into the CubeMLP module. Results: We evaluate our model on three benchmark datasets: CMU-MOSI, CMU-MOSEI, and AVEC2019. Experimental results demonstrate that our method achieves a CCC of 0.594 and an MAE of 4.23 on AVEC 2019, outperforming existing baselines by a CCC improvement of 0.483 and an MAE reduction of 2.14. On CMU-MOSEI, our model achieves an Acc2 of 86.2%, showcasing strong generalization across domains. Discussion: These results verify the effectiveness and robustness of our approach for practical applications in multimodal affective computing and mental health assessment. The proposed framework addresses key limitations in cross-modal alignment and feature fusion while demonstrating superior performance over existing methods. Conclusion: Our study presents a novel multimodal depression detection framework that leverages semantic enhancement, graph-based fusion, and optimized deep learning components. The improved performance across multiple datasets highlights its potential for real-world applications in mental health diagnostics and affective computing.

Read the paper · More papers on PaperTik