Deep Learning Based Video Captioning in Bengali
Amir Hossain Raj, Ashek Seum, Aurpan Dash, Saiful Islam, Faisal Muhammad Shah · 2021
Generating meaningful textual descriptions from visual contents having the context in consideration is very challenging in terms of Natural Language Processing (NLP) and Computer Vision (CV) and video captioning adds more to it. Currently, the field of video captioning is highly explored in different languages except for Bengali, where no work is available so far. In this paper, we have proposed a model to generate meaningful captions describing the activities of a video in Bengali. The proposed model is an encoder-decoder-based novel deep architecture that incorporates a combination of different Convolutional Neural Networks (CNNs), like 2D-CNN and 3D-CNN with Bidirectional Long Short Term Memory (Bi-LSTM) as the encoder and two-layer Long Short Term Memory (LSTM) as the decoder. Lack of proper concern in this field left us with no video dataset with proper Bengali captions. For this reason, all the English captions of the MSVD dataset have been translated to Bengali language using the Google Translate API. In this work, we have trained and evaluated our model on the MSVD dataset and achieved a 32.6%, and 51.2% score on BLEU and CIDEr, respectively, which is currently the state-of-the-art result even compared with closely related Bengali image captioning works.