Multimodal Image-Text Representation Learning for Sketch-Less Facial Image Retrieval

Dawei Dai, Yingge Liu, Shiyu Fu, Guoyin Wang · 2024

Sketch-less facial image retrieval (SLFIR) framework aims to break the barriers that drawing a high-quality facial sketch requires excellent skills and substantial time, it performs the retrieval using a partial sketch with as few strokes as possible. However, such early-stage sketches often contain only local details, resulting in poor retrieval performance. In this study, we propose learning of the representation by fusing the sketches with prior human semantic knowledge to improve the early retrieval performance. Specifically, (1) based on the LAION-Face dataset, a facial language-image pretraining (FLIP) model is constructed to learn the aligned representations of facial image and text; (2) subsequently, using FLIP as the backbone, multiscale features of sketch and text are extracted and fused to learn the efficient representation for the final retrieval. The proposed method achieves state-of-the-art early retrieval performance on all two public datasets and exhibits a good generalization ability in practical testing.

Read the paper · More papers on PaperTik