Enhanced Dense Image Captioning Based On Transformers
Tilottama Goswami, Sathvika Potu, Kuntluri Prasanna Reddy, Veltoor Prashamsa, K. Ram Mohan Rao, Mukesh Kumar Tripathi · 2024
The paper introduces a pioneering work that explores the fusion of computer vision and natural language processing for narrative generation. We propose an innovative methodology that combines the GRiT model for dense captioning with the GPT model for story generation. GRiT extracts detailed object descriptions from images, while GPT constructs cohesive storylines based on these descriptions. The integrated approach aims to generate narratives with visual and textual information. Through experimental validation and qualitative analysis, we demonstrate the effectiveness of our method in creating engaging stories from visual content. Our paper advances AI-driven narrative generation and opens avenues for applications in digital storytelling, content creation, and creative AI.