FD-CSS: Nested Cross-Modal Fusion for Survival Prediction
Mengke Li, Lanlan Chen, Chenyang Wu · 2025
Aiming at the problems of noise interference, mode conflict and performance limitation of static fusion methods in medical multimodal data fusion, this paper proposes a cancer diagnosis and survival prediction model FD-CSS based on dynamic multimodal fusion. ResNet-50 and sparse graph convolutional network (SGCN) were combined to extract high-order features from pathological images and gene expression data, respectively, and a nested dynamic fusion framework was designed. At the feature level, information entropy encoder was used to evaluate the importance of features, and sparse gating strategy was used to suppress noise. At the modal level, the classification confidence is quantified based on the true class probability (TCP), and the multi-modal representation is fused with adaptive weighting. Based on the BLCA, BRCA and KIRC datasets of TCGA database, the superiority of FD-CSS in cancer subtype classification (AUC) and survival analysis (C-index) tasks was verified. The results showed that the AUC (0.864) and C-index (0.691) of FD-CSS in the BRCA dataset were improved by 1.6% and 4.2%, respectively, compared with the second-best method, and the standard deviation across the datasets was the lowest (such as the AUC standard deviation of KIRC ±0.016), indicating its strong generalization ability. Ablation experiments further demonstrated the core contribution of the dynamic fusion module (C-index of BRCA decreased by 24.0% after removal) and the effectiveness of ResNet and SGCN in modeling image and genetic features (AUC increased by 15.7% in BLCA and C-index increased by 1.9% in BRCA). This study provides a dynamic fusion framework for multimodal medical diagnosis that takes into account both classification and survival analysis, and provides a new idea for high robustness modeling of clinical decision support systems.