Image Caption Generation using Contrastive Language Image Pretraining

G Bharathi Mohan, R Harigaran, P Sri Varshan, Repakula Srimani, R Prasanna Kumar, R Elakkiya · 2024

Image captioning is a problem within the boundary of both natural language processing and computer vision. A challenging task is making well-written sentences. It needs both understanding language rules and meaning. Describing picture contents with accurate sentences impacts individuals with low vision to understand images better. In this study, we have introduced a novel approach of using Contrastive Language–Image Pre-training(CLIP) encodings as image features for an Lstmbased textual decoder model trained on the coco dataset 2017. On training with textual context CLIP model has rich image semantic features. For datasets with data in large-scale and diverse features, CLIP provided meaningful captions without any additional attention mechanism or pre-training because of joint embedding learning. To evaluate the model we have generated captions for random unseen images. The model performed well with the unseen images generating meaningful captions related to the image.

Read the paper · More papers on PaperTik