Vision Transformers Explained Series
Skylar Callis, USDOE National Nuclear Security Administration (NNSA). Office of Defense Programs (DP) · 2024
Since their introduction in 2017 with Attention is All You Need¹, transformers have established themselves as the state of the art for natural language processing (NLP). In 2021, An Image is Worth 16x16 Words successfully adapted transformers for computer vision tasks. Since then, numerous transformer-based architectures have been proposed for computer vision. This article walks through the Vision Transformer (ViT) as laid out in An Image is Worth 16x16 Words.