Boosting Long-Tailed Recognition With Label Descriptor and Beyond
Zhengzhuo Xu, Ruikang Liu, Zenghao Chai, Qi Yan, Lei Li, Haiqin Yang, Chun Yuan · IEEE Transactions on Multimedia · 2025
Long-Tailed Recognition (LTR) poses significant challenges due to the heavily imbalanced nature of real-world data, which severely skews data-driven deep neural networks. Despite the rapid progress of Vision-Language Models (VLMs), they still face challenges in effectively learning from long-tailed visual data. In this paper, we present a comprehensive analysis of the reasons behind the underperformance of VLMs and propose a hierarchical inference framework to address this issue. Specifically, we prompt the large language models to generatesentence-leveldescriptors for class labels and conduct the open vocabulary classification by computing the average similarity between the image and each descriptor. Areweightingmechanism is further proposed to filter out uninformative descriptors. To mitigate model bias incurred by the long-tail distribution, we propose a feature adapter with the logit adjustment technique and fine-tune the CLIP model via visual prompt tokens. We introduce the Shared Feature space Mixup (SFM) to enhance the interaction between modalities to address tail visual feature insufficiency. Finally, we propose a hierarchical inference manner to combine the aforementioned proposals. Extensive evaluations demonstrate that our approach achieves state-of-the-art performance by fine-tuning only a few parameters on the Places-LT, ImageNet-LT, and iNaturalist 2018 benchmarks.