Multi-level Fusion in a Hybrid Architecture for 3D Image Segmentation
Zhiyuan Li, Zuguo Chen, Hejun Huang, Chaoyang Chen · 2024
In this paper, multi-level feature fusion is applied in a hybrid CNN-Transformer architecture. In more detail, level-by-level fusions of multiple levels of Transformer output with various scales of CNN encoder feature maps are used in the decoder. The Transformer will be used to capture long-distance dependencies, while the CNN will be fully utilized for local feature extraction. It is also possible to combine features of various sizes to create discriminative features, which aids the decoder step in producing finer segmentation results. The method MFTr proposed in this paper has been validated with the BraTS2020 dataset for brain tumor segmentation task. The experimental results show the superior performance of MFTr in segmenting 3D medical images.