Enhancing Image Representation for Caption Generation Using Graph Convolutional Networks
P.S Pavan, Rupa Sarkar, Deeplata Sharma · 2024
Graph convolutional networks (GCNs) have emerged as an effective approach for photograph illustration and caption-era obligations. In this paper, we gift a singular GCN architecture for photo caption technology that models visible relationships in a photo. The proposed architecture consists of essential additives: a kernel-based photograph illustration based totally on graph convolutional networks (GCNs) and a caption generator primarily based on the generated picture representation. To generate image captions, the generated photo illustration is fed to the caption generator, which is skilled in the use of supervised brand new. We compare our method on commonly used datasets that require the ability to generate captions for pix and demonstrate that our architecture achieves overall performance. Our outcomes show that by exploiting the nearby structure of modern-day images, the use of GCNs can successfully improve photograph captioning overall performance and imply the path of state-of-the-art future research in this field.