Self-attention as the backbone: A survey on Vision Transformers

Ahmad Waseem, Pietro Ruiu, Seth Nixon, Andrea Lagorio, Mássimo Tistarelli · Computer Vision and Image Understanding · 2026

The human cognitive system efficiently processes complex environments by focusing on salient regions, inspiring attention mechanisms that amplify critical information. Deep learning models integrate attention to enhance performance across tasks, with self-attention gaining prominence through its remarkable success in Transformer-based natural language processing. This success propelled its adoption in computer vision, where Transformer-based models surpass conventional deep networks on multiple benchmarks. Recently, large-scale pretraining has enabled vision foundation models, where self-attention acts as a scalable backbone for learning transferable visual representations across diverse tasks. By modeling global dependencies without heavy inductive biases, self-attention supports flexible and expressive representation learning. However, Vision Transformers (ViTs) rely on large parameter counts and high computational budgets, and their quadratic complexity with respect to input size restricts scalability and real-world deployment. Consequently, substantial research has focused on reducing computational overhead while preserving or improving performance through self-attention optimization. In parallel, State Space Models (SSMs) have emerged as an alternative paradigm, replacing attention with linear-time recurrence for efficient long-range modeling. This survey provides a comprehensive review of advanced backbone techniques that improve ViTs through novel self-attention designs or hybrid integrations, mainly for image classification, including their transfer to multimodal large language models, generative frameworks, and emerging SSM-based architectures, alongside ViT explainability approaches. We first outline self-attention and Transformers, then introduce a taxonomy of self-attention variants. Finally, we systematically compare attention architectures in terms of design, accuracy, and computational efficiency, and discuss current limitations and future research directions.

Read the paper · More papers on PaperTik