Self-Attention with Spatial Decay Matrix Based on Inverse Multiquadric

Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Jehwan Choi, Kang-Hyun Jo · 2025

Vision Transformers have achieved great performance across computer vision and foundation model tasks. At the core of the vision transformer, self-attention learns spatial interactions across visual tokens, exhibiting weak inductive biases. Although self-attention tends to capture long-range dependencies from input tokens, the positional information between tokens is underexplored. Recent methods exploited the absolute and relative positional embedding that supplements the input token and attention matrix. This paper presents an alternative method, adding spatial priors to the attention matrix based on the distance from a query to other tokens in a decay way. Neighborhood regions located around a target query have greater attention in their spatial matrix, and regions far from a target query receive lower attention in their attention scores. This design allows each query to attend to all tokens while perceiving various spatial interactions in different regions. Additionally, this paper reduces the spatial redundancy of the vision Transformer based on a hierarchical design. Experiments are trained and evaluated on the benchmark dataset ImageNet-1K image classification. As a result, the proposed method surpasses vision Transformer DeiT-T by 0.3% Top-1 accuracy and reduces the computational cost of DeiT-T by 75% GFLOPs. This verifies that this paper improves the efficiency of global self-attention while still keeping the high accuracy of vision Transformer.

Read the paper · More papers on PaperTik