Improved Cross-Modal Retrieval Systems Using Self-Reinforcement and Quadruplet Alignment Hashing
Xiaoqing Liu, Zhiwen Yu, Jun Jiang, Bin Wang, Fa Zhu, Xingchi Chen, Witold Pedrycz · IEEE Transactions on Consumer Electronics · 2025
Cross-modal retrieval presents significant challenges for consumer technology applications, demanding innovative approaches to bridge semantic gaps between different data modalities while ensuring efficient information access. This paper introduces a novel Self-Reinforcement and Quadruplet Alignment Hashing (SRQA) framework specifically designed to enhance cross-modal retrieval systems for improved user experiences. Our approach distinguishes itself through three key contributions. First, we develop a dynamic unified similarity matrix that adaptively balances label-driven semantic information with modality-specific correlations, enabling more nuanced cross-modal representations than traditional fixed alignment strategies. Second, we propose a novel quadruplet-based hashing method that implements an efficient hard sample mining strategy through the refinement of both absolute and relative distance constraints between samples, thereby providing a more precise and efficient semantic alignment mechanism for cross-modal retrieval. Third, through extensive experiments conducted on three benchmark datasets—MIRFLICKR-25K, NUS-WIDE, and MS-COCO—our framework consistently outperforms ten state-of-the-art cross-modal retrieval methods across various hash code lengths, offering significant advancements for consumer technology applications requiring efficient multi-modal information retrieval.