Bridging Visual Representation and Efficiency for Resource-Constrained Devices
Adityam Ghosh, Xinfeng Ye, Sathiamoorthy Manoharan · 2024
Transformer models have transformed computer vision and natural language processing, setting new standards of performance on benchmark datasets. However, these models are not optimized for resource-constrained devices like mobile and edge devices. This paper introduces CLIP-Mobile, an efficient approach for enhancing mobile visual representation learning by aligning image features with textual annotations. Inspired by CLIP-Lite, CLIP-Mobile requires only a single negative image-text pair per positive pair during contrastive learning, significantly reducing the required batch size compared to CLIP. CLIP-Mobile employs a mobile-centric architecture with MobileNet as the image encoder and the lightweight Lite Transformer for the text encoder. Evaluations were carried out to measure CLIP-Mobile's efficacy. Despite being pretrained on only 10% of the MS-COCO dataset, it achieves zero-shot top-5 accuracy close to a CLIP model trained on the same subset of data. The findings highlight CLIP-Mobile's potential as a robust and resource-efficient framework for advancing visual representation learning for mobile and edge devices, enabling real-world applications in constrained computational environments.