Energy-Based Video Captioning With Distance Learning
Thanh Trung Le, Hiep Xuan Huynh · 2025
We propose a novel framework combining EnergyBased Models (EBMs) and Energy Distance to address the video captioning problem. This framework is designed to bridge the semantic gap between visual and linguistic data while learning effective multimodal representations. By leveraging EBMs, the model evaluates the compatibility between video features and linguistic descriptions and optimizes this compatibility through Energy Distance. Experimental results on the benchmark MSRVTT dataset demonstrate the framework's effectiveness, with consistent reduction in energy scores and the ability to generate coherent and contextually appropriate captions, particularly for shorter videos. However, the method also faces challenges in handling diverse scenarios and longer videos, suggesting potential for further improvement. These results validate the framework's feasibility and highlight future research directions to enhance video captioning performance in real-world applications.