Image captioning based on low-level grid features, segmentation features, and high-level fusion features
Yuxin Li, Qiuming Zhang · 2025
Existing image captioning algorithms primarily rely on visual features and generated partial descriptions to predict subsequent words. However, these visual features often lack contextual or detailed object information, leading to captions that may not accurately describe the content of the image. To address these issues, we propose an image captioning approach that integrates low-level grid features, segmentation features, and high-level fusion features. This method supplements visual information with segmentation features, incorporating them alongside grid features in the encoder. To enhance the model's ability to capture visual context, we introduce memory-augmented attention within the existing IILN module. In the decoding phase, we leverage simultaneous utilization of low-level grid features, segmentation features, and high-level visual semantic fusion features to comprehensively learn multi-level representations of image regions and semantic relationships. Experimental results demonstrate that our approach generates more precise descriptions, achieving a competitive CIDEr score of 136.5 on the MS COCO "Karpathy" offline test split.