γ -CRD:Gamma-Cooperative Retrieval Diffusion Model for Robust Incomplete Multimodal Learning

Ruiting Dai, Wenwei Zhu, Zheyu Wang, Haoran Meng, Zhengdao Yuan, Yandong Yan, Lisi Mo · 2025

Multimodal learning in open environments faces significant challenges due to modality incompleteness and noise interference. Current prompt engineering emphasizes modality absence over task-instance contextualization that hinders cross-modal knowledge transfer, whereas conditional generation approaches for missing modality recovery exhibit an excessive reliance on the quality of available modalities. To address these issues, we propose a Gamma-Cooperative Retrieval Diffusion model (γ-CRD), inspired by the human brain's multi-source contextual completion mechanism, which leverages a retrieval-augmented prompt generation framework and normal-inverse Gamma noise modeling to enhance robustness in incomplete multimodal learning. Specifically, it consists of three modules: (1) Retrieval-Augmented Contextualization: Construct a multimodal memory bank and retrieve relevant instances via similarity calculation under a gating mechanism to augment the contextualization of missing modalities. (2) Prompt-Driven Diffusion Generation Module: Builds prompts based on retrieval results and incorporates them into a denoising diffusion probabilistic model through an attention mechanism to enhance contextualized knowledge transfer and generate missing modalities. (3) Inverse-Gamma Noise Optimization Module: Model a mixed normal-inverse gamma distribution, which is dynamically aware of noise and enables uncertainty estimation in multimodal fusion, ensuring robust and reliable multimodal regression. Extensive experiments on three real-world datasets demonstrate that γ-CRD consistently outperforms state-of-the-art baselines, especially achieving a 5.7% improvement in accuracy compared to the leading model e.g., IMDer [1].

Read the paper · More papers on PaperTik