Utilizing Adaptive Feature Scaling and Dynamic Contextual Graph Networks to Improve Captioning of Images Through the Recognition of Unseen Relationships

Yugandhara A. Thakare, Kishor H. Walse, Mohammad Atique · Apple Academic Press eBooks · 2026

Automated systems for image captioning are increasingly important in fields like assistive technology, content management, and robotics. However, current systems often fall short when it comes to accurately capturing unseen relationships and the deeper contextual meaning in images, which limits their practical use. These systems typically depend on convolutional neural networks (CNNs) and basic attention mechanisms, which result in captions that lack semantic depth and lead to subpar performance in recognizing both objects and their relationships. This paper presents a cutting-edge approach to image captioning that combines adaptive feature scaling for enhanced object detection, dynamic contextual graph networks for improved comprehension of contextual information, meta-learning for 254 zero-shot detection to identify previously undiscovered relationships, and a multimodal attention mechanism to produce more detailed and precise captions. This framework demonstrates superior performance over existing models, particularly in the classification of objects and their relationships, making it highly applicable for automated social media content generation and image analysis. During testing, the model showed a 4.9% boost in object classification precision and a 4.5% increase in accuracy, along with a remarkable improvement in relationship recognition, achieving an 8.3% increase in precision and an 8.5% improvement in accuracy.

Read the paper · More papers on PaperTik