Advancing Medical Image Registration with the Vision Foundation Model (VFM): A Modular Pre-trained Framework

Haojie Wang, Qingying Zhou, Zhifang Pan · 2024

In this paper, we introduce a novel pretraining registration framework that enhances medical image registration by leveraging the Visual Foundation Model (VFM). Our frame-work is designed to concurrently learn registration and refine the segmentation capabilities of the VFM, with SAM-Med3D chosen for its superior performance due to extensive training on diverse 3D medical datasets. The architecture integrates a segmentation network based on VFM and a registration network, where multi-channel feature fusion and a multi-head cross-task attention mechanism are employed to improve spatial information extraction and prevent overfitting. The Dual-Adapter mechanism is also introduced, fine-tuning feature representations of both moving and fixed images at each encoder layer, thereby enhancing segmentation accuracy and alignment precision. Our experiments, conducted on the Learn2Reg Challenge and The Cancer Imaging Archive (TCIA) datasets, demonstrate the effectiveness of our framework. Significant improvements in performance metrics, including Dice coefficient and HD95, were observed across multiple state-of-the-art deep learning-based registration methods when integrated with our framework. Ablation studies further validate the contributions of the multi-head cross-task attention mechanism and the multi-channel feature fusion module, showing notable declines in performance when these components are removed. Our framework significantly outperforms traditional registration methods and advanced segmentation-assisted models, highlighting its robustness, generalizability, and enhanced understanding of anatomical information across diverse medical imaging modalities.

Read the paper · More papers on PaperTik