Improving CLIP Model Efficiency: A Zero-Shot Approach Using Multiple Smaller Models
Giang Truong Le, Nhan Thanh Tran, Nam Mai Hoai Tran, Vinh Dinh Nguyen · 2025
In the rapidly evolving landscape of video content, efficient retrieval systems are essential to manage the vast amounts of data available.This research introduces a zero-shot method that leverages an ensemble of smaller CLIP models to enhance retrieval performance.By embedding image-text pairs into a FAISS database , we facilitate efficient similarity searches and normalize confidence scores across different models for consistent comparison.Our method combines OpenAI ViT-L-14-336 and Apple ViT-L-14 models , demonstrating that their aggregated outputs can match or surpass the performance of larger models like Apple ViT-H-14 and META ViT-H-14.The proposed system not only improves retrieval accuracy but also optimizes memory usage, making it suitable for real-time applications and resource-constrained environments.