Enhancing Video Captioning: A Bayesian Normalized Attention-based Multi-Dimensional Graph Network with Moss Growth Optimization
Khallikkunaisa, H. S. Niranjana Murthy, Shirish Kulkarni, Syed Mohd Uzair Iqbal, Umang Pancha, Natrayan L · 2025
The explosive increase of video data requires the creation of advanced automated video captioning systems that can produce effective textual descriptions. Current methods, such as traditional deep learning (DL) models, are plagued with drawbacks such as high computational expense, overfitting, and inefficacy in contextual modeling, resulting in poor caption generation. For resolution of these difficulties, this contribution introduces the Bayesian Normalized Attention-based Multi-Dimensional Graph Network with Moss Growth Optimization (BN-AMGN-MGO) model, engineered to improve video captioning automatization on the MSVD set. The proposed framework starts off with pre-processing, in which unnecessary frames are removed via template matching to get only the representative frames. The feature encoding process utilizes a Bayesian Normalized Neural Network (BNNN) for dynamically manipulating probabilistic weight distributions to counter overfitting and increase robustness. During the decoder process, a Hierarchical Attention-based Multi-Dimensional Edge Graph Neural Network (HAM-GNN) combines temporal context modeling, graph neural networks, and hierarchical attention mechanisms to create structured textual descriptions. Classification stage employs a temporal attention mechanism to project dialogue act labels into attention-weighted representations, utilizing speaker dependencies for better accuracy. To further optimize performance, MGO is employed to refine HAM-GNN’s hyperparameters through wind direction determination, spore dispersal search, dual propagation search, and cryptobiosis mechanisms, ensuring effective exploration and exploitation of the search space. The proposed BN-AMGN-MGO model achieves the lowest MSE (0.021), highest PSNR (34.89), and SSIM (0.905), ensuring high-quality caption generation. Additionally, it outperforms existing models with an accuracy of 94.5% while maintaining the shortest computational time (5.6 sec), making it a robust and efficient solution for hierarchical attention-based video captioning.