A Transformer-based Multimodal Feature Fusion Model for Video Captioning

Yudong Wang · 2024

Video Captioning requires effective extraction and fusion of multimodal features, including visual, semantic, and textual information, to generate accurate natural language descriptions. To enhance semantic representation and detail capture in complex scenes, a multimodal feature fusion-based Video Captioning model is proposed, integrating global visual features, local visual features, and video semantic features. The model employs a Hybrid Attention Mechanism (HAM) and a multi-head attention fusion module, using the SlowFast model for visual feature extraction, Detectron2 for local feature recognition, and the CLIP model for semantic feature extraction. Experimental results on the ActivityNet dataset demonstrate the model's ability to generate coherent Video Captioning in multi-object and fine-grained action scenarios. Although the model lags behind state-of-the-art methods in BLEU-4 and CIDEr metrics, it shows significant potential for improvement in feature fusion and attention mechanisms, offering strong support for video understanding and automatic description generation in complex environments.

Read the paper · More papers on PaperTik