TUTNet: Leveraging transformers for comprehensive extraction and preservation of global information in medical image segmentation

Minquan Zhao, Jiahe Yu, Hui Qi, Qin Gu, Miao Wang, Yaduan Ruan · Biomedical Signal Processing and Control · 2025

In the field of medical image segmentation, the introduction of Transformers to assist in medical image segmentation has proven to be effective in recent years. Compared to traditional CNN methods, Transformers have excellent global context information extraction capabilities. However, using a purely Transformer-based architecture for image segmentation not only increases the overall network parameters but also decreases the ability to extract local features. The accurate extraction and integration of local and global features are crucial for achieving medical image segmentation. Therefore, we propose TUTNet, an improved medical image segmentation network. TUTNet features a dual encoder architecture, consisting of a CNN-based encoder and a Transformer-based encoder. The former effectively extracts local information, while the latter captures global information. Skip connections simultaneously receive information from both encoders, eliminating semantic differences between the two encoders through a cross-attention mechanism. Skip connections are also of high importance for medical image segmentation. In our network, the skip connections utilize a purely Transformer mechanism, fully leveraging the information extracted by encoders at all levels, utilizing a self-attention mechanism to focus on channel information, and then employing a cross-attention mechanism to query and filter global information with local information, thus extracting valuable information for decoders at each level. We conducted extensive validation on four datasets, demonstrating the effectiveness of our network across different feature datasets. Our work can be viewed on https://github.com/whycantChinese/TUTNet . • Using a CNN-Transformer dual encoder to achieve the simultaneous extraction of global and local information. • Proposing a special spatial cross-attention mechanism that enables the elements of the context matrix to extract spatial information. • Applying multi-head self-attention and the proposed spatial cross-attention mechanism in skip connections.

Read the paper · More papers on PaperTik