AAPAR: CLIP-Based Adaptation and Alignment for Pedestrian Attribute Recognition
Li Xu, Hefei Ling, Yuxuan Shi, Zongyi Li, Ping Li · IEEE Transactions on Biometrics Behavior and Identity Science · 2025
Pedestrian Attribute Recognition (PAR) aims to identify attributes of target pedestrians in images. While current methods improve performance through attribute localization and correlation modeling, CLIP-based approaches demonstrate superior effectiveness by leveraging textual semantics to mine attribute relationships. However, two critical limitations persist: (1) existing CLIP adaptation methods in feature extraction networks fail to adequately bridge the gap between CLIPs design purpose (single-label modality alignment) and PARs task nature (multi-label recognition with inter-attribute correlations); and (2) inadequate exploitation of attribute hierarchies. To address these challenges, we propose AAPAR, a CLIP-based adaptation and alignment model with two key contributions: (1) a modalityaligned CLIP adaptation method that integrates Transformer layers as adaptation modules after the image encoder while preserving crucial pre-adaptation image-attribute caption alignment, jointly enhancing multi-label dependencies while preserving original alignment knowledge; and (2) a hierarchical classification architecture comprising global-context, part-aware local, and refined individual attribute classifiers that collectively model attribute relationships at multiple granularities. Extensive experiments on standard PAR datasets (PETA, PA-100K, and RAPv2) and UPAR2024 dataset demonstrate that AAPAR outperforms state-of-the-art methods. The source code will be released on https://github.com/lixusky/AAPAR.