Multidimensional Semantic Augmented Visual Storytelling

Bing Li, Guangheng Jia, Xiyan Gao, Can Ma · 2024

Visual storytelling is a multi-modal generation task of crafting a consistent narrative for a set of sequential images. The challenge of visual storytelling is that it requires the model not only to identify the specific content of individual image, but also to connect all the elements to form a consistent and complete story. Previous methods employ scene graphs constructed by external knowledge to associate various images. However, ignoring global features may cause these methods to generate incoherent content. Additionally, there is an accuracy problem in the construction of scene graphs. The incorrect relationships between objects in the scene graph can introduce noise to the model's learning process and inject an inaccurate bias term during inference. Inspired by human storytelling, we simulate human cognition and propose a Multidimensional Semantic Augmented Network, by introducing various textual semantic information to deal with the multi-modal generation task. Concretely, we first propose a Scene Semantic Extraction Module, which takes advantage of scene textual semantic information as the central theme to ensure that each sentence is coherently related to the overall narrative. Moreover, we introduce an Object-and-Action Semantic Augmented Module to enrich the details of images from the object-and-action textual semantic perspective, which further bridges the semantic gap between visual and language modalities. Finally, extensive experiments are conducted on the popular visual storytelling dataset VIST, and the results demonstrate that our proposed method achieves competitive performance on several auto-metrics.

Read the paper · More papers on PaperTik