The Storyteller: Computer Vision Driven Context and Content Generation System
Anwar ul Haque, Sayeed Ghani · Research Square · 2022
Abstract The human capability of detecting, understanding, and contextualizing objects in the real world by machines has always been a dream for computer scientists. Along with other important and pending challenges in the field of computer vision, image captioning with context and content is an important research problem. In our research, we attempted to develop a human-like storytelling system that can caption images with the perspective of content, context, syntax, and knowledge. Our methodology is a combination of Capsule Networks for image encoding, Knowledge Graph for content and context awareness, and Transformer Neural Networks for decoding. During feature extraction, spatial, geometrical, and orientational details are extracted using Capsule Networks. To equip our model with content, context, and semantics, the corpus is passed through the Knowledge Graph. The decoding phase is a combination of Knowledge Graph and Transformer Neural Network for knowledge-driven captioning. Dynamic multi-headed attention in the decoder is used for memory optimization. Our model is trained over MSCOCO. The results provide good content and context understanding along with B4: 18.23, M: 19.2, R: 41.1, and C: 54.19. The primary outcome of our research is generating autonomous story-type captions for real-world images.