LP²CR-IoT: Lightweight and Privacy-Preserving Cross-Modal Retrieval in IoT
Mingyue Li, Yuting Zhu, Ruizhong Du, Chunfu Jia · IEEE Internet of Things Journal · 2025
As a pivotal link between visual and linguistic relationships, image-text cross-modal retrieval has received widespread attention. However, existing studies primarily focus on intricate machine learning models to enhance retrieval accuracy and overlook the critical aspect of privacy preservation for images and texts, rendering them unsuitable for lightweight IoT environments. To tackle these challenges, we propose LPCR-IoT, a lightweight and privacy-preserving cross-modal retrieval scheme tailored to IoT environments. LPCR-IoT employs knowledge distillation to train lightweight student models for extracting feature vectors from images and texts, subsequently embedding them into a unified semantic space. Significantly, we propose a new training metric (i.e., Intra-modal Consistent Contrast Loss), which improves the retrieval accuracy by increasing the semantic consistency of the image and text in the common embedding space. Additionally, a novel quadtree index structure leveraging hybrid representation vectors is designed to effectively mitigate retrieval overhead, where feature vectors of images and texts alongside representation vectors are encrypted using a secure kNN algorithm based on LWE, enabling image-text matching in a large-scale ciphertext environment. Finally, we provide a detailed formal analysis to evaluate the security of LPCR-IoT and validate its practicality through extensive experiments on three real-world datasets, namely COCO, Flickr30k and NUS-WIDE.