AI-Generated Image Captions Evaluation: In Terms of Semantic Similarity and Conciseness
Suguru Tsujioka, Kojiro Watanabe, Akihiro Tsukamoto, Muneyuki Natsume · 2024
The increasing trend of visual content on social media requires the best output from an image captioning system, which can describe the main content of any given image. The current study undertakes the performance evaluation of two AI models, LLaVA and BLIP, in generating appropriate image captions that are comparable to those generated by humans. We utilize WordNet to measure semantic similarity, which tests generated AI captions against human references, with emphasis on the conciseness and relevance of the same. Our findings suggest that LLaVA tends to perform better than BLIP in producing summary-like captions that are closer to human-like descriptions. The evaluation methodology proposed herein is robust yet allows for semantic variations while the identification of the main subjects representing images is appropriately captured. This work will help improve more human-like AI captioning systems, which are crucial in marketing and social media analysis applications.