Advancing Video Captioning via Visual-Linguistic Feature Fusion
Guanghua Chen, Bin Fang · International Journal of Pattern Recognition and Artificial Intelligence · 2025
Video captioning is a challenging multimodal task that combines computer vision and natural language processing. Previous methods primarily extract useful information from visual features by designing sophisticated encoders. Natural language descriptions generated solely from visual features often fall short of expected performance due to the semantic gap between the visual modality and the textual modality. In this paper, we propose a novel encoder–decoder-based approach for video captioning, which enhances video feature representations by incorporating object and action-centric linguistic features from upstream encoders. Specially, we leverage a cross-attention mechanism to incorporate textual information into visual features, thereby enhancing their expressive capability. Additionally, we introduce a fusion layer to facilitate the interaction between heterogeneous representations and mitigating irrelevant noise. Extensive experiments on the MSVD, MSR-VTT and VATEX datasets demonstrate the superiority of our method and achieve significantly superior performance across standard evaluation metrics.