Generalized Energy Distance for Video Captioning: A Novel Framework
Thanh Trung Le, Cong Thang Pham, Hiep Xuan Huynh · 2025
The rapid growth of video content on digital platforms necessitates intelligent systems for generating natural language descriptions of videos, a task that demands effective alignment of spatiotemporal visual features with linguistic representations.Existing video captioning models often face challenges in achieving semantic alignment, resulting in suboptimal captions.To address this, we propose a novel video captioning framework based on Generalized Energy Distance (GED), incorporating a ResNet-101 Video Encoder, a multihead attention Caption Encoder, and a Transformer-based Decoder with cross-modal attention.Evaluations on the MSR-VTT dataset demonstrate state-of-the-art performance, with optimal 𝛼 values (1.0, 1.5) balancing loss minimization and caption quality, highlighting the importance of hyperparameter tuning for robust and coherent caption generation.