Content-Guided Video Captioning With Cross-Modal Transformer for Driving Scenario Understanding

Zhibo Zhou, Chenhao Cui, Yang Yang, Boyang Wang, Xianghong Tang, Norah Saleh Alghamdi, Zhoujun Li, Feiran Huang · IEEE Transactions on Intelligent Transportation Systems · 2025

Video captioning constitutes a challenging multi-modal learning task, particularly within Advanced Driving Assistance Systems (ADAS), with the objective of automatically generating precise and diverse descriptive sentences for real-time understanding of driving scenarios. Previous methods have predominantly concentrated on independent subtasks such as detection, tracking, and segmentation, which constrain their capability to comprehensively describe autonomous driving events. However, these approaches frequently neglect the vast collective knowledge embedded in human-labeled sentences, particularly regarding domain-specific vocabulary in traffic scenarios. To address this limitation, we propose a Content-Guided Video Captioning (CGVC) model for driving scenario understanding to identify the pivotal content in videos through a supervised learning approach. Specifically, we introduce a cross-modal Transformer to effectively capture the correlation between appearance and motion features for the video captioning task. Additionally, a content-guided module is devised to predict salient content information from videos by leveraging the collective expressions found in human-labeled sentences. The CGVC model enhances caption generation performance by transferring the traffic knowledge learned from the content-guided module to the video captioning task. We conduct extensive experiments on the MSR-VTT, MSVD and TVC datasets, demonstrating that our model outperforms other competitive methods.

Read the paper · More papers on PaperTik