Bridging Crowd Analysis and Natural Language: A New Vision Language Approach for Behavioral Video Captioning
Selen Gürbüz, Turan Göktuğ Altundoğan, Mehmet Karaköse · IEEE Access · 2026
Crowd analysis is a critical application area for addressing challenges in domains such as transportation and public safety within smart cities. Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have enabled the integration of computer vision tasks with natural language understanding, opening new opportunities for crowd behavior analysis. This study aims to analyze the behavior of crowded human groups through natural language descriptions. However, a major limitation in this domain is the lack of suitable video captioning datasets, despite the availability of image-text alignment datasets for crowd analysis. To address this gap, we construct a novel dataset by generating synthetic and semantically rich captions using the Gemini Flash API. The caption generation process leverages ontological features derived from ground truth annotations and keyframes extracted from public crowd video datasets. The resulting dataset is used to train multiple VLM architectures, enabling the generation of detailed behavioral descriptions from crowd videos. Experimental results demonstrate strong performance, achieving 91.81 BERTScore, 42.23 ROUGE-L, and 72.9 CIDEr. These results highlight the effectiveness of the pro-posed approach in bridging the gap between visual crowd analysis and natural language understanding.