Multi-level spatio-relational segformer (MLSRS-SegFormer): A novel vision transformer with adaptive spatial induction and dynamic positional encoding

Inda Rusdia Sofiani, Hadi Suyono, Erni Yudaningtyas, Fitri Utaminingrum · MethodsX · 2025

Medical image segmentation is foundational to precision medicine. However, state-of-the-art Vision Transformers (ViTs) inherently suffer from a critical trade-off between comprehensive global contextualization and robust local boundary discrimination, especially in high-variance clinical data. This deficit necessitates a novel architecture. This study introduces the Multi-Level Spatio-Relational SegFormer (MLSRS-SegFormer), a novel vision transformer architecture designed to significantly enhance semantic segmentation through adaptive spatial induction strategies, dynamic positional encoding, and refined local context learning. Our proposed Multi-Level Spatio-Relational SegFormer (MLSRS-SegFormer) model demonstrates significant architectural innovation, superior performance in comparative experiments, and robust validation for clinical applications, as summarized in the following key points: • MLSRS-SegFormer integrates three clear and novel contributions beyond standard SegFormer: (1) Adaptive Patch Weighting in PatchEmbedding for dynamic feature induction, (2) Hausdorff-bias Attention for explicit spatial prioritization, and (3) Relative Positional Encoding (RPE) for nuanced and adaptive spatial relationship understanding. • Comparative experiments reveal MLSRS-SegFormer's superior performance with consistent gains in segmentation accuracy, achieving the highest mIoU (0.968) and mDSC (0.980). Crucially for clinical applications, the model also demonstrates the lowest HD95 (1.1668) , which validates its exceptional boundary precision . • Bland-Altman analyses further confirm its near-zero systematic bias and remarkable consistency in area and boundary delineation, providing robust and highly accurate segmentation vital for clinical applications despite a longer inference time.

Read the paper · More papers on PaperTik