Multivariate Feedback-Based Image-Text Joint Learning for Sketch-Less Facial Image Retrieval

Yingge Liu, Dawei Dai, Guoyin Wang, Shuyin Xia · IEEE Transactions on Circuits and Systems for Video Technology · 2025

Sketch-Less Facial Image Retrieval (SLFIR) framework facilitates the retrieval of target images with minimal strokes through a human-computer interactive approach, thereby circumventing the need for high-quality sketches required by traditional frameworks. The primary approach utilizes a contrastive learning framework that minimizes the distance between sketch images and their target images in the embedding space, while maximizing the distance from non-target images, thus efficiently learning representations of sketches and images. However, during the initial stages of sketching, the sparse strokes that capture only partial facial features can inadvertently match non-target facial images, blurring the distinctions between positive and negative samples and impairing early retrieval performance. To overcome this challenge, we introduce a multimodal retrieval model based on diversified feedback reinforcement learning, which not only enhances the semantic integrity of sketches but also optimally ranks the sketches corresponding to positive samples using diversified feedback. Specifically, (1) we developed a Facial Language-Image Pre-training (FLIP) model and, leveraging this model, constructed an on-the-fly multimodal retrieval model that excels in recognizing sparse and exaggerated sketches by extracting and integrating multiscale features from both sketches and textual descriptions. (2) Furthermore, we implemented a novel reward mechanism that adjusts the rewards for target images, accommodating reasonable fluctuations in sketch rankings on actual images. This mechanism effectively differentiates similar images during retrieval, ensuring a more consistent and progressively improving ranking list. Extensive experiments validate that our proposed method significantly enhances early retrieval accuracy and generalization capability.

Read the paper · More papers on PaperTik