Fine-Grained Visual Classification: A Survey of Methods, Architectures, and Emerging Trends
Md Al Amin, Ho Yub Jung · IEEE Access · 2026
Fine-grained visual classification aims to distinguish visually similar subcategories within the same semantic class, which remains challenging due to subtle inter-class differences and large intra-class variation. This survey presents a structured review of fine-grained visual classification methods published between 2015 and 2026, emphasizing learning paradigms and architectural design principles rather than performance alone. We introduce a taxonomy that categorizes existing approaches into supervised, lightweight, augmentation-based, multimodal and interpretable, self-supervised methods, bilinear CNNs, destruction and construction learning, and analyze representative techniques within each category by abstracting common processing pipelines and design choices.We further summarize widely used benchmark datasets and evaluation protocols, provide an extensive quantitative comparison across multiple datasets, and analyze trends in backbone architectures, learning paradigms, dataset usage, and performance variability. Our analysis shows that supervised attention-based and transformer-driven models currently achieve the strongest performance, while lightweight, augmentation-based, self-supervised, and multimodal methods provide complementary advantages in efficiency, scalability, and interpretability. In addition, Bilinear CNNs enhance fine-grained discrimination through second-order feature interactions, whereas Destruction and Construction Learning (DCL) improves part-aware representation by perturbing and reconstructing image structures to emphasize discriminative local details. Finally, we discuss open challenges and highlight promising directions for future research in fine-grained visual classification.