Validating Image Captioning Models Using Text-to-Image Algorithms via Generative AI

Kunal Lall, Anant Lall · 2025

The potential of employing text-to-image models as validation tools for assessing the output of image captioning algorithms is investigated in this work. Using 5,000 images from the COCO dataset and textual descriptions produced by eight distinct image captioning models, we produce visuals. This work uses CLIPScore, a metric from the multimodal CLIP model, to evaluate the similarity between original and regenerated images using five well-known open-source text-to-image models. The findings show that CLIPScore offers useful information about the relative quality of captioning models, even though it is not entirely consistent with the evaluation of visual quality. Furthermore, tests with different prompt formats demonstrate that images with extremely thorough explanations are more like the originals than those with minimal descriptions. However, the scope was constrained by limited computer capabilities, resulting in results that were based on a smaller sample of the information. Nevertheless, the methodology provides a starting point for future studies on generative AI-based image captioning validation.

Read the paper · More papers on PaperTik