Cross-Modal Hash Retrieval Based on Deep Learning

Fanqi Meng, Kaidi Tian, Jingdong Wang, Yan Tao Tian · 2024

The exponential growth of multimedia data increasingly requires retrieval technology across different data modalities, such as searching for videos through images and searching for sounds through text. This cross-modal retrieval technology aims to explore the intrinsic semantic connections between different modalities, so that information of one modality can be retrieved from data of another modality. However, due to the difficulty of feature extraction and the high cost of data annotation, cross-modal learning is difficult and the retrieval efficiency is low. Deep learning methods have demonstrated powerful capabilities in both data processing and image processing, bringing new possibilities to solving the above cross-modal retrieval problems. Hashing methods have also been widely used in multimodal retrieval technology due to their advantages such as low storage cost and fast retrieval speed. This article summarizes the main requirements of cross modal retrieval tasks, reviews some mainstream research methods based on deep learning for cross modal hash retrieval, and finally lists some existing problems and possible future research trends in the field of cross modal retrieval, which will provide meaningful references for future researchers.

Read the paper · More papers on PaperTik