Med-T: a pixel-level position information aware transformer for medical visual task
Jianbo Huang, Ke Zhao, Jing Wang, Kexin Zhang · 2024
The model of transformer structure occupies a dominant position in the field of multimodal large model. While previous studies have highlighted the potential of Visual Transformer (ViT) models, their reliance on large datasets poses challenges in domains like medicine, where obtaining extensive data can be difficult. In such scenarios, traditional convolutional neural networks (CNNs) often outperform transformer-based models due to their ability to capture pixel level fine-grained information. In this paper, we proposed soft mask operation and fine-grained information awared visual transformer Med-T, a CNN-Transformer hybrid visual backbone network tailored for visual feature extraction task on limited datasets. Through extensive evaluation across three small datasets, Med-T consistently out performs alternative approaches, showcasing the efficacy of leveraging the pixel-level position information extraction ability of CNN branch.