Deep Cross-modal Hashing Retrieval Based on Semantics Preserving and Vision Transformer

Jinlin Hong, Huayong Liu · 2022

In response to the problem of similarity measure differences in different similarity coefficients that occur in cross-modal multi-label retrieval, this article uses an interval parameter to correct this bias. A new supervised hash method is proposed by introducing the transformer structure which performs well in CV and NLP tasks into cross-modal hash retrieval, called the Deep Semantics Preserving Vision Transformer Hashing (DSPVTH). This method uses network structures such as vision transformer to map different modal data into binary hash codes. It also uses the similarity relationship of multiple tags to maintain the semantic association between different modal data. Validation on four commonly used multimodal text datasets, Mirflickr25k, NUS-WIDE, COCO2014 and IAPR TC-12, shows a 2% to 8% improvement in average accuracy compared with the current optimal method, which means our method is robust and effective.

Read the paper · More papers on PaperTik