Saliency Feature Learning for Multimodality Person Retrieval in Visual Internet of Things
Shuai You, Yujian Feng, Shitao Wang, Fei Wu, Yuchen Sha, Yimu Ji, Xiao‐Yuan Jing · IEEE Internet of Things Journal · 2025
In the visual Internet of Things (VIoT), smart surveillance is an important component of the multimodality person retrieval (MPR) task. Capturing discriminative pedestrian information from images aids in identifying target individuals through sketch or text descriptions, improving the reliability of VIoT systems. Current granularity-level matching methods face challenges with information redundancy and modality differences in MPR tasks. In this article, we propose a novel approach for intra and intermodality saliency feature learning by multigranularity feature selection and relation (MFSR). First, to reduce redundant information within each modality, the intramodality feature selection module (IFSM) employs the adaptive weighting mechanism to enhance salient pedestrian features while suppressing irrelevant features. Second, the multigranularity feature relation module (MFRM) aligns cross-modality salient person features by increasing the similarity scores of local and global features across modalities, to reduce differences between multimodality visual-text pairs. Finally, the cross-modality similarity matching (CSM) loss is designed to enhance consistency in visual-text pairs of the same identity, ensuring compact intraclass features by minimizing discrepancies between cross-modality similarity distributions and identity-matching distributions. Experimental results show that our approach achieves state-of-the-art performance on benchmark datasets.