Image Caption Generator with CLIP Interrogator

Adhiraj Ajaykumar Jagtap, Jitendra Chandrakant Musale, Sambhaji Nawale, Parth Takate, Indra Kale, Saurabh Waghmare · 2025

This research paper presents a cutting-edge system for creating image captions. It combines the strengths of the COCO (Common Objects in Context) model with the CLIP (Contrastive Language-Image Pre-training) Interrogator. Our method allows us to generate both short social media captions and in-depth descriptive captions for a broad range of images. By bringing together COCO's skill in spotting objects and CLIP's deeper understanding, the system produces captions that fit the context. We then use a Natural Language Processing (NLP) technique to polish these captions making sure they're easy to read and make sense. This two-pronged approach to caption creation makes content more accessible and boosts the impact of social media posts. As a result, the system can adapt to many different uses, from creating content to improving digital accessibility. This groundbreaking method creates fresh opportunities to interact between humans and AI when it comes to understanding and describing visual content. It sets the stage for image captioning systems that are more natural and aware of context down the road.

Read the paper · More papers on PaperTik