U-BERT for Fast and Scalable Text-Image Retrieval

Tan Yu, Hongliang Fei, Ping Li · 2022

Exploiting cross-modal attention on image region features and text features, cross-modal BERT models have achieved higher accuracy than the embedding-based methods in cross-modal text-image retrieval. Nevertheless, cross-modal BERT models take image-text pairs as input, requiring a quadratic computational complexity. Thus, cross-modal BERT models are prohibitively slow and not scalable. A remedy is a two-stage strategy, wherein the first stage uses an embedding-based method to retrieve top K items and the second stage deploys the heavy cross-modal BERT to re-rank these K items. Nevertheless, to achieve a satisfying accuracy, K should be large, making the retrieval in the second phase still slow. In this paper, we propose a U-BERT model to achieve an effective and efficient cross-modal retrieval. The proposed U-BERT decomposes each image/text feature into an intra-modal component and an inter-modal component. In the first stage, U-BERT only uses the intra-modal component of the image/text features to obtain the text-image similarity scores based on two independent encoders, with a linear computation complexity. In the second stage, U-BERT reuses the intra-modal component and additionally use the inter-modal component as the complementary residual to enhance the intra-modal component's discriminating capability. Benefited from reusing the intra-modal component, we only need few Transformer layers to generate the inter-modal component, making the second phase efficient even facing a large K. Extensive experiments on public benchmarks demonstrate the efficiency and effectiveness of the proposed U-BERT.

Read the paper · More papers on PaperTik