CLAIR: Evaluating Image Captions with Large Language Models

David M. Chan, Suzanne Petryk, Joseph E. Gonzalez, Trevor J. Darrell, John Canny · 2023

The evaluation of machine-generated image captions poses an interesting yet persistent challenge.Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, caption diversity, and specificity.Existing highly-engineered measures attempt to capture specific aspects, but fall short in providing a holistic score that aligns closely with human judgments.Here, we propose CLAIR 1 , a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) to evaluate candidate captions.In our evaluations, CLAIR demonstrates a stronger correlation with human judgments of caption quality compared to existing measures.Notably, on Flickr8K-Expert, CLAIR achieves relative correlation improvements over SPICE of 39.6% and over image-augmented methods such as RefCLIP-S of 18.3%.Moreover, CLAIR provides noisily interpretable results by allowing the language model to identify the underlying reasoning behind its assigned score.

Read the paper · More papers on PaperTik