TA-MIR: Text Aggregation for Multimodal Feature Representation in Medical Image Registration

Yuhe Dai, Zhiyong Huang, Weimin Huang · 2025

The aggregation of multimodal features in medical image registration remains underexplored, limiting the performance of current models in capturing complex anatomical relationships. Traditional convolutional neural networks (CNNs) often overlook the rich semantic information available from text, while existing approaches lack effective methods to combine spatial and contextual cues. In this paper, we propose Text Aggregation for Medical Image Registration (TA-MIR), a novel framework that enhances encoder-decoder architecture by incorporating anatomical text embeddings throughout the registration process. By employing large kernel blocks for improved receptive fields in U-Net and fusion blocks at each level, our model effectively integrates image features with semantic text information. Extensive experiments on three brain MRI datasets-OASIS, IXI, and LPBA40-demonstrate that our approach achieves state-of-the-art performance, significantly improving registration accuracy and anatomical coherence compared to traditional CNN and Transformer-based methods.

Read the paper · More papers on PaperTik