A Unified Generative Hashing for Cross-Modal Retrieval
Junfeng Tu, Xueliang Liu, Yanbin Hao, Richang Hong · ACM Transactions on Multimedia Computing Communications and Applications · 2025
Cross-modal hashing is a highly effective and efficient method for information retrieval, enabling the search for correlated data across different modality databases using compact hash codes. Conventional cross-modal hashing typically uses separate model structures for each modality and aligns approximate continuous representations of the final hash codes. These approaches not only require specialized models for each modality but also introduce a gap between the discrete hash codes and their continuous features, yielding only approximate alignment. To address these issues, we propose a unified generative cross-modal hashing method that leverages a single Uniform Mixture-of-Expert Decoder (UMoED) for both image and text modalities. UMoED streamlines cross-modal hash learning by integrating two key design elements: (1) a cross-modal representation unification module that employs unified queries to consolidate modality-specific features into a common space, and (2) an adaptive expert enhancement module that adaptively enhances feature modeling based on the input modality. Furthermore, our decoder-based hashing method outputs hash codes in a generative manner, producing precise representations of discrete codes to bridge the gap between the discrete and continuous space, thus ensuring precise alignment during similarity learning. Extensive experiments on three benchmark datasets demonstrate that the proposed method achieves the state-of-the-art performance in cross-modal hashing retrieval.