Enhancing Kidney Tumor Segmentation in MRI Using Multi-Modal Medical Images With Transformers

Srisopitsawat Pavarut, Joong-Sun Lee, Takashi Obi, Masaki Kobayashi, Hajime Tanaka, Yoh Matsuoka, Yasuhisa Fuji · IEEE Access · 2025

Detecting and segmenting kidney tumors is a significant challenge due to their small size, heterogenous appearance, and complex characteristics. Clinicians typically rely on various diagnostic techniques, including multi-modal medical imaging, to achieve accurate diagnoses. In deep learning, however, conventional Convolutional Neural Network (CNN)-based approaches face limitations in handling incomplete or missing modality data. In this study, we propose a novel multi-modal fusion framework built on the transformer architecture, which captures both local and global context through self-attention mechanisms. Our approach integrates Contrast-Enhanced Computed Tomography (CECT) and MRI scans, including multiple phases and sequences, to enhance segmentation accuracy and improve robustness over existing methods. A key innovation of our framework is a parameter-free multi-modal attention mechanism that dynamically weighs the contribution of available modalities, enabling the model to handle missing data while selectively focusing on the most informative sources. This design not only improves segmentation performance but also enhances interpretability for clinical use. We further conduct a comprehensive comparison with conventional multi-head attention mechanisms, demonstrating the efficiency and effectiveness of our approach. Experimental results show that our method consistently outperforms single-modality baselines and existing fusion strategies, particularly in challenging cases such as Chromophobe RCC, while requiring fewer parameters and lower computational cost. Overall, the proposed framework demonstrates strong potential for advancing kidney tumor segmentation by offering a flexible, efficient, and clinically applicable solution under realistic multi-modal data conditions.

Read the paper · More papers on PaperTik