RetViT: Retentive Vision Transformers

Shreyas Dongre, Shrushti Mehta · 2024

Attention mechanisms have traditionally served as the dominant approach for computing weighted relationships between elements in input tokens for both Natural Language Processing (NLP) and vision tasks, however a significant paradigm shift has emerged within the domain of NLP. This shift introduces “Retention” as a novel and computationally efficient alternative, consistently delivering equivalent or superior results. This research applies pure Retention directly to sequences of image patches, replacing two-dimensional attention blocks from Vision Transformers (ViT) for image classification tasks along with accelerated model training and inference speeds. Proposed model integrates retention into Compact Convolutional Transformer (CCT) to build upon CCT’s impressive $84.1 \%$ Top1-acc on ImageNet-1k, utilizing a mere 4 million parameters. Post-retention our model surpasses other models with significantly high parameters in accuracy, achieving a state-of-the-art $91.57 \%$ accuracy on ImageNet-1k with only 6.5 million parameters and only 3.65G FLOPs. Implementation code and pretrained weights for the model will be made publicly available.

Read the paper · More papers on PaperTik