Enhancing Interpretability of NesT Model Using NesT‐Shapley and Feature‐Weight‐Augmentation Method
Li Xu, Lei Li, Xiaohong Cong, Huijie Song · IET Computer Vision · 2025
ABSTRACT The transformer's capabilities in natural language processing and computer vision are impressive, but interpretability is crucial in specific domain applications. The NesT model, with its pyramidal structure, demonstrates high accuracy and faster training speeds. Unlike other models, a unique aspect of NesT is its avoidance of the [CLS] token, which presents challenges when applying interpretability methods that rely on the model's internal structure. Instead, NesT divides the image into 16 blocks and processes them using 16 independent vision transformers. We propose the NesT‐Shapley method, which utilises this structure to combine the Shapley value method (a self‐interpretable approach) with the independently operating vision transformers within NesT, significantly reducing computational complexity. On the other hand, we introduced the feature weight augmentation (FWA) method to address the challenges of weight adjustment in the final interpretability results produced by interpretability methods without [CLS] token, markedly enhancing the performance of interpretability methods and providing a better understanding of the information flow during the prediction process in the NesT model. We conducted perturbation experiments on the NesT model using the ImageNet and CIFAR‐100 datasets and segmentation experiments on the ImageNet‐Segmentation dataset, achieving impressive experimental results.