Leaner and Faster: Two-Stage Model Compression for Lightweight Text-Image Retrieval
Siyu Ren, Kenny Qili Zhu · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies · 2022
Current text-image approaches (e.g., CLIP) typically adopt dual-encoder architecture using pre-trained vision-language representation.However, these models still pose non-trivial memory requirements and substantial incremental indexing time, which makes them less practical on mobile devices.In this paper, we present an effective two-stage framework to compress large pre-trained dual-encoder for lightweight text-image retrieval.The resulting model is smaller (39% of the original), faster (1.6x/2.9xfor processing image/text respectively), yet performs on par with or better than the original full model on Flickr30K and MSCOCO benchmarks.We also opensource an accompanying realistic mobile image search application.