SAMPL: Self-Attention Modelled Patch Learning for Efficient Visual Understanding
Zhiming Hu, Salar Hosseini Khorasgani, Weiming Ren, Iqbal Mohomed · 2025
We study the patch selection problem for efficient transformer-based visual understanding wherein a sampler can be used to drop less informative patches at inference in order to speed up model execution on resource-constrained devices. As no labels are available on the saliency of the patches, existing works either try to solve an auxiliary task of locating distinctive patches or learn a policy network through the global image/video-level supervision. The former approach could drop redundant but important patches while the latter suffers from the weak supervision of a single class label per image/video. In this work, we observe that the attention weights in trained transformer-based models clearly highlight the salient regions in images and videos. Therefore, we propose a learned patch sampling framework called SAMPL that utilizes the attention weights as fine-grained patch-level supervision to learn a lightweight policy network for patch selection. To train SAMPL end-to-end with the transformer-based models, we introduce a new loss function based on the REINFORCE algorithm to match the distribution of patch selection probabilities and the attention scores. Experimental results on ImageNet, UCF101, Something-Something v2 and Kinetics-400 show that SAMPL can effectively increase the throughput by at least 1.5× while achieving competitive classification accuracy.