Leaner and Faster: Two-Stage Model Compression for Lightweight Text-Image Retrieval

Siyu Ren, Kenny Qili Zhu · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies · 2022

Current text-image approaches (e.g., CLIP) typically adopt dual-encoder architecture using pre-trained vision-language representation.However, these models still pose non-trivial memory requirements and substantial incremental indexing time, which makes them less practical on mobile devices.In this paper, we present an effective two-stage framework to compress large pre-trained dual-encoder for lightweight text-image retrieval.The resulting model is smaller (39% of the original), faster (1.6x/2.9xfor processing image/text respectively), yet performs on par with or better than the original full model on Flickr30K and MSCOCO benchmarks.We also opensource an accompanying realistic mobile image search application.

Read the paper · More papers on PaperTik