ROP: Exploring Image Captioning with Retrieval-Augmented and Object Detection of Prompt

Meng Zhang, Kai Yang, Shouqiang Liu · 2024

Image captioning models aim to bridge the modalities of vision and language by generating natural language descriptions that match the content of an input image. Existing methods for generating image captions do so by integrating visual encoders with language models, and by training the model using large volumes of image-text pairs. However, this approach leads to a significant increase in the model's parameter size due to the need to store extensive visual concepts and detailed textual descriptions. With substantial investments in datasets and computational resources, the quality of image caption generation has notably improved, though these advancements come at a high cost. In this paper, we introduce a novel method, termed ROP, that combines retrieval and detection to address this challenge. This approach significantly reduces the number of trainable parameters while preserving the accuracy of the descriptions. Through the retrieval module, the model can find the k most similar sentences to the input image, utilizing this rich contextual information to enhance the overall understanding of the image content. Simultaneously, the detection module enhances the model's interpretation of fine details by identifying prominent regions within the image. Our method has proven its effectiveness and feasibility in experiments.

Read the paper · More papers on PaperTik