Image caption generation via improved vision-language pre-training model: perception towards image retrieval

Roshni Padate, Ashutosh Gupta, Mukesh Kalla, Arvind Sharma · The Imaging Science Journal · 2025

A novel approach using an improved VLP is proposed which operates in two phases utilizing Flickr 8k, Flickr 30k and COCO datasets. In Phase 1, relevant features are extracted from Flickr 8k images and indexed using an enhanced Chi-square test. Phase 2 comprises offline and online processes. The improved VLP model generates query text from Flickr 30k images during the offline process. In the online process, query text undergoes feature extraction similar to Phase 1. Extracted features are then subjected to an improved Fisher score-based ranking process and stored alongside indexed features for similarity checks in a database. By integrating VLP and sophisticated feature extraction methods, the approach enhances image retrieval by providing more nuanced and accurate insights into image content. Finally, the improved VLP acquired the highest Bleu-1 score of 0.896 for dataset 1 at the training percentage of 80, while conventional methods achieved lower ratings.

Read the paper · More papers on PaperTik