Tagsim: Topic-Informed Attention Guided Similarity Metric for Image Caption Comparison

Vipul Chanchlani, Vishal Himmatsinghka, Ayush Himmatsinghka, Jivnesh Sandhan, Tushar Sandhan · 2025

Existing image caption evaluation metrics, such as BLEU, ROUGE, and CIDEr primarily rely on high-level similarities like n-gram matching. Here, we propose TAGSim, a novel metric that automatically incorporates topics or concepts as the caption’s topic contains the main summary of the image. TAGSim creates topic-weighted latent representations, thereby acquiring semantics by integrating both con-text and topic information. It leverages a regression model to create an attention-based novel similarity metric, while internally building caption representations. TAGSim moves be-yond lexical overlap by focusing on meaningful semantic relationships, better aligning captions with the core image topic. It handles paraphrasing and diverse expressions, ensuring a more nuanced and reliable evaluation across languages and styles. It outperforms the strong similarity baseline by an average 0.72 points (SCI evaluation index) across 3 datasets. It is language agnostic, empirically established that it is a pseudometric, and correlates well with standard caption evaluation metrics. Our code and datasets will be publicly available.

Read the paper · More papers on PaperTik