Hybrid Attention Vision Transformer-based Deep Learning Model for Video Caption Generation
Khustar Ansari, Priyanka Srivastava · 2025
Video caption generation is an extremely challenging computer vision task that automatically generates the caption for video clips with a clear understanding of the embedded semantics. Due to the extensive use of images for conveying large volumes of information, there exists a huge demand for image analytics in information processing systems. Various techniques are designed for automatic image caption generation to solve the computer vision challenges, but highlighting the contextual information more accurately still results in a challenging task. To overcome such issues, a Hybrid Attention Vision Transformer-based Convolutional Neural Network-Gated Recurrent Unit-Bidirectional Long Short-Term Memory (HAVisT-CNN-GRU-BiLSTM) is developed for video caption generation. Specifically, the Vision transformer (VisT) effectively captures the global dependencies and learns the intricate relationships present in the image. Moreover, the proposed approach combines the salient information extracted from the CNN, GRU, and BiLSTM models resulting in less memory usage, high dimensional data processing, and improved convergence speed. Moreover, the hybrid attention-based Vision transformer facilitates capturing the global context information and enhances computational efficiency. The experimental results demonstrate that the HAVisT-CNN-GRU-BiLSTM model achieves the efficient metric values of BLEU, METEOR, and ROUGE as 0.48, 0.28, and 0.71 respectively outperforming the other existing techniques.