FTAN: Exploring Frame-Text Attention for lightweight Video Captioning
Zijian Zeng, Yali Li, Yue Zheng, Siqi Li, Shengjin Wang · 2023
Traditional video captioning approaches employ LSTM as a lightweight decoder. However, these methods focus on fully extracting visual features, but pay less attention to textual information, resulting in relatively low-quality performance. Recent transformer-based methods achieve more accurate results, but at the cost of excessive computing resources. In this paper, we propose a lightweight model for video captioning named Frame-Text Attention Network (FTAN), aiming to make full use of both visual and textual features to obtain more accurate captions. We develop a novel text attention module in FTAN, which uses the hidden state of LSTM as query to generate attentive text features. Then the attentive text features are merged with visual features, which are used as input for LSTM to generate more accurate captions. To the best of our knowledge, we are the first to introduce attention mechanism to extract more textual information hidden in LSTM architecture in video captioning. Extensive experiments demonstrate the effectiveness of FTAN. FTAN outperforms the state-of-the-art LSTM-based method on MSVD dataset by 0.8 in CIDEr-D and is about one-fourth of the transformer-based methods in terms of parameters.