Cross-Modal Uncertainty Modeling With Diffusion-Based Refinement for Text-Based Person Retrieval
Shenshen Li, Xing Xu, Chen He, Fumin Shen, Yang Yang, Heng Tao Shen · IEEE Transactions on Circuits and Systems for Video Technology · 2024
Text-based person retrieval (TBPR) is a challenging task that aims at retrieving candidate pedestrian images from a gallery, using textual descriptions as queries. Existing methods generally assume that the textual query and the unique candidate image have a certain cross-modal relationship under one-to-one constraint, and optimize their conditional probability via a discriminative paradigm. However, in real scenarios of TBPR, a textual query may associate with multiple candidate images at one time, indicating that the uncertainty resides in the one-to-many cross-modal relationship. Moreover, the learnt conditional probabilities from the discriminative paradigm of existing methods may be less effective in reflecting the joint probabilities of the textual query and candidate images. To tackle these problems, we propose a novel method termed Cross-modal Uncertainty Modeling with Diffusion-based Refinement (CUMDR) for the TBPR task. First, we implicitly model the cross-modal uncertainty to capture richer semantics and complex correlations, thus generating diverse yet plausible retrievals. Additionally, to reasonably mitigate the impact of noisy data with high uncertainty, we quantify the uncertainty to allocate the importance of raw and complement annotations, which are generated from the multi-modal large language model based on the retrieval-augmented template. Finally, we propose a novel diffusion-based denoiser to progressively refine cross-modal alignments by learning joint probabilities. Extensive experiments on three TBPR datasets demonstrate the superior performance and generalizability of our CUMDR approach compared to the latest methods. Our anonymous implementation repository is available athttps://github.com/Shenshen7/CUMDR.