Enhanced image captioning with advanced context-aware object relational model
Madhvi Patel, Pranay Deepak Reddy Vaka, Dhirendra Pratap Singh, Jaytrilok Choudhary, Surendra Solanki · Discover Computing · 2025
Image captioning generates text in natural language to describe a given image. Recent advances in various object detection with attention mechanisms pushed for exploring multiple image captioning methodologies to create more meaningful and accurate captioning models. Although existing pipelines do well in describing the image, there has not been enough emphasis on relationship modeling between image features. Relationship modeling is crucial for establishing context between various objects in an image. We have proposed a novel Advanced Context-Aware Object Relational Model (ACAORM) that not only improves relationship modeling between image features but also better sentence generation due to Transformer architecture. ACAORM builds relation-aware visual representations for image description and builds captions with prior attention to relevant Regions of Interest (RoI). We tested the proposed methodology with three widely used datasets, Flickr8K, Flickr30K, and MS-COCO. The results show that it beats numerous cutting-edge techniques. ACAORM scores 0.3526, 0.4439, and 0.8813 in $$ BLEU-4 $$ on Flickr8k, Flickr30k and MS-COCO, respectively, outperforming popular cutting-edge models such as VitaCap on Flickr8k and AGF on Flickr30k and MS-COCO by 10.88%, 47.96% and 139.48% on the respective benchmarks.