A Multi-Stage Vision Transformer for Fine-grained Image Classification
Zheng Huang, Ji‐Xiang Du, Hongbo Zhang · 2021
Different from general image classification, Fine-grained visual classification(FGVC) aims at recognizing the subclass of an object. This task is challenging due to the differences between subcategories that might be very small. At present, Vision Transformer(ViT) achieved huge success in traditional classification tasks benefited from its innate attention mechanism. Recent works extend the vision trans-former into the FGVC task and achieve state-of-the-art performance. However, those transformer-based structures inevitably harm the discriminative regions while also requiring a high computational cost. To address these problems, we propose a pooling-based vision transformer model with overlapping patches(OP-ViT), which inserts the pooling layer into a transformer encoder to form a multi-stage hierarchical structure. In particular, the overlapping patches scheme can preserve the informational region effectively. We evaluate the model on two fine-grained mushroom datasets. The experimental results show that the proposed model achieves competitive performance with fewer parameters and computational overhead.