Prompt-Guided Transformer and MLLM Interactive Learning for Text-Based Pedestrian Search

Zefeng Lu, Ronghao Lin, Yap‐Peng Tan, Haifeng Hu · IEEE Transactions on Information Forensics and Security · 2025

Aiming to retrieve pedestrian images based on a textual description query, Text-Based Pedestrian Search (TBPS) gains increasingly attention due to its applications in security surveillance. As a fine-grained classification task, TBPS requires identifying images of individuals with different semantic contexts yet the same identity, as well as distinguishing images of individuals who share similar appearances but distinct identities. Consequently, TBPS is challenged by semantic variations in positive pairs and appearance similarity between negative pairs. To tackle these challenges, we propose the Prompt-guided Transformer and MLLM Interactive learning (PTMI) model to learn identity-discriminative representations across different modalities. PTMI consists of three components: the Prompt-guided Transformer (Promformer), MLLM Interactive Learning (MIL) and Dual-branch Cross-modal Learning (DCL). Firstly, the Promformer is designed to handle semantic variations in positive pairs by introducing learnable prompts, composing of three types: instance-shared, instance-specific and layer-specific. Optimized by Cross-modal Intra-class Consistency (CIC) loss, these prompts minimize intra-class variations and retrieve positive images with various semantics. Secondly, the MIL component is introduced to address appearance similarity between negative pairs by focusing on key image patches and description words filtering by the local discriminator. Powered by Multimodal Large Language Model (MLLM), the local discriminator adopts soft attention to highlight important image regions and descriptive words, which preserves semantic information while emphasize discriminative details. Lastly, the DCL integrates global and local branches to bridge modality discrepancies. The global branch employs SDM loss for heterogeneous distribution alignment, while the local branch applies Anchor-Based Contrastive (ABC) loss for instance-level contrastive learning. Unlike conventional contrastive loss, ABC loss leverages MLLM features as anchors to decouple modality and semantic differences, enhancing alignment efficiency. Extensive experiments on three TBPS datasets have validated the effectiveness of PTMI.

Read the paper · More papers on PaperTik