Integrating Region Proposals with Recurrent Neural Networks for Image Paragraph Captioning
Yogita Ambure, Riya Hulule, Sakshi Paste, Samriddhi Datir, Sharmila Kharat · 2025
Capturing rich and coherent paragraphs describing images, this represents an abstract "image paragraph captioning," a challenging task. The area marks a timely intersection of computer vision and natural language processing. In this work, we present an improved framework for the generative creation of paragraph captions for images. This research gives the brief of the method that exploits Region Proposal Networks (RPNs) and Convolutional Neural Networks (CNN), encoder and multilevel decoder using RNN, ensuring a more effective paragraph caption generation process.We evaluated our method on the Stanford Image Paragraph dataset and presented its ability to generate semantically rich and coherent paragraphs. Experimental results demonstrate that the presented model achieves competitive performance, while improving significantly caption diversity and coherence compared to other state-of-the-art approaches.