MMPL-Seg: Prompt Learning with Missing Modalities for Medical Segmentation

Ruowen Qu, Wenxuan Wu, Yeyi Guan, Lin Shu · 2024

Using text prompts has been proven helpful in improving medical segmentation performance. However, text-prompted segmentation of the pulmonary infectious areas is a challenging task, for two major reasons: (i) the infected area is prone to misjudgment, resulting in under-segmentation or over-segmentation, (ii) full medical prompts are expensive, especially during real-world during testing, so missing medical note modality often occurs. To address these issues, we propose a new method for Prompt Learning with Missing Modalities for Medical Segmentation (MMPL-Seg). Specifically, based on the UNet-like encoder-decoder framework, we first introduce a Modality Reconstruction and Fusion module (MRF), which is suitable for both complete and incomplete modalities. Based on joint training of two branches of image-only and image-text pairs, this module enables image-only samples to also have latent semantic information in the test phase. In addition, we use the Dynamic Convolution Layer (DyConv) to achieve dynamic fusion of deep image-text features into shallow edge layers. To validate our idea, we conduct a series of experiments on the QaTa-COV19 multi-modality dataset. Experiments on the QaTa-COV19 dataset demonstrated that our model outperforms state-of-the-art approaches under complete and missing modality scenarios, proving its high performance and robustness.

Read the paper · More papers on PaperTik