Generating Spatially-Aware Dense Video Captions for Indoor Human Behavior Analysis with Position-Based Scene Knowledge

Bin Chen, Yugo Nakamura, Shogo Fukushima, Yutaka Arakawa · 2024

Current dense video captioning methods do not effectively capture spatial information about the individual and the surrounding objects in the scene.This limitation can lead to ambiguous captions, making it difficult to detect abnormalities in human activities in the human-centric applications, such as security surveillance and remote indoor caregiving, where detecting abnormalities in human activities is crucial and these activities are closely intertwined with the spatial information of the objects in the scene.We propose a system that enhances dense video captioning of indoor human behavior by integrating two key sources of information: spatial context extracted by Retrieval-Augmented Generation (RAG) from a knowledge base that contains the positions and categories of all objects in the room, and spatial information from RGB-D pairs that includes the positions and categories of people and objects within the Field of View (FOV).By combining these elements, our approach reduces ambiguity and improves the accuracy of behavior descriptions.Our proposed baseline method is evaluated on the custom dataset, achieving scores of 15.7 for METEOR, 20.3 for ROUGE-L, 52.3 for Recall and 52.4 for Precision.These results demonstrates that our system effectively addresses the limitations of PDVC and GVL in collecting spatial information and provides improvements over SG-PDVC and SI-PDVC in this area.We propose a convenient and effective method for generating spatial information-enhanced captions to describe human behavior in untrimmed videos.

Read the paper · More papers on PaperTik