E2E-LANet: Unleashing the Potential of Linear Attention for High-Resolution Medical Image Segmentation via Pretraining Distillation
Hui Tang, Gang Yu, Lu Cao · 2024
Medical image segmentation plays an important role in computer-aid diagnosis. In the past years, convolutional neural networks, especially the UNet-based architectures with symmetric U-shape encoder-decoder structure and skip connection, have been widely used in a variety of medical imaging segmentation tasks. However, such methods are difficult to capture global context. In contrast, Vision Transformers (ViTs) utilize the self-attention mechanism to obtain long-range dependencies in images. Despite these advantages, the huge computational costs of the self-attention mechanism reduce the inference speed of models, thereby limiting their real-world applications. Meanwhile, hierarchical pixel-level decoding of U-shape architectures increases the complexity of large-size segmentation decoding. To address these limitations, in this paper we propose a linear attention-based end-to-end model for high-resolution medical image segmentation, named E2E-LANet. It consists of two modules: an image encoder based on constructed Efficient Vision Transformer (EVT) that reduces the high quadratic complexity of self-attention to linear complexity and a Query-Supported Mask Decoder (QSMD) that employs independent mask query for each type of medical images in segmentation mask predictions. It focuses on the detailed segmentation decoding of patch-level features to further improve the efficiency of large-size mask predictions. Especially, we leverage the advantages of pretraining distillation to improve the transferability of our model in various segmentation tasks. Experimental results on the nucleus segmentation task and the polyp segmentation task display the important role of pretraining distillation in improving the performance of our framework. E2E-LANet demonstrates superior performance compared to CNNs, ViTs, and latest Mamba-based U-shape architectures.