Crossmodal knowledge distillation hashing

Jinhui Wang, Lu Jin, Zechao Li, Jinhui Tang · Scientia Sinica Technologica · 2021

Hashing is widely used for large-scale crossmodal retrieval due to its high computation and storage efficiency. Most of the existing crossmodal hashing methods separately generate hash codes for single-modal data, ignoring the context information within and between modalities. Therefore, the learned hash codes cannot preserve the potential dependence between multimodal data. To address this issue, this paper proposes a crossmodal hashing method based on knowledge distillation. The Transformer-based teacher network is first employed to capture the inter- and intra-modal context information from images and texts to further learn the joint representations that can contain rich visual semantic-associated information. The teacher network then generates discriminate hash codes by projecting the joint representations into a compact hamming space. In addition, knowledge distillation is leveraged to transfer the potentially associated knowledge embedded in the teacher network to the student network, such that the hash codes generated by the student network can preserve the associated information of the multimodal data as much as possible. The proposed method is evaluated on the MIRFLICKR-25K, NUS-WIDE, and MS-COCO datasets. The experimental results show that the performance of the proposed method is better than that of the other compared methods, demonstrating its superiority.

Read the paper · More papers on PaperTik