Multi-view–enhanced modal fusion hashing for Unsupervised cross-modal retrieval
Longfei Ma, Honggang Zhao, Zheng Rong Jiang, Mingyong Li · 2023
Cross-modal hashing is an important direction for multimodal data management and applications, which has recently received more and more attention. Unsupervised cross-modal retrieval does not rely on tag information and is more applicable to the real world. However it still faces some problems. Existing methods mainly encode for local features or global features. Due to the effect of negative samples, it is easy to cause noise interference. To solve these problems, we propose a Multi-view–enhanced modal fusion hashing for Unsupervised cross-modal retrieval (MUCH) to improve these problems. Firstly, we propose a multi-view network. Images inherently contain richer semantics, and we employ a multi-view network to observe the image from different perspectives and obtain the overall and local features of the image. Secondly, we introduce a noise cancellation module to approximate the cross-modal data feature alignment from both intra-modal and cross-modal perspectives before generating the hash code. Finally, we construct a distribution-based similarity weighting matrix to replace the graphical similarity matrix. And we performed multi-view enhancement experiments on JDSH and CIRH, with 1% to 2% enhancement over DAEH on all three datasets.