Human Action Captioning based on a GRU+LSTM+Attention Model

Lijuan Marissa Zhou, Weicong Zhang, Xiaojie Qian · 2021

To quickly understand human actions in the videos, this paper proposes to solve the human action captioning problem which aims to automatically generate text descriptions based on human action videos. A sequence-to-sequence method based on GRU+LSTM+Attention (GLA) model is proposed to solve this problem. Specifically, GRU is applied as the encoder to capture the temporal information of actions. The LSTM is applied as the decoder to generate the fluent fine-grained descriptions for human actions. To focus on the most relevant part of actions and capture the correlation between actions and descriptions, an attention mechanism is applied in the proposed method. Experiments on the WorkoutUOW-18 dataset demonstrate the effectiveness of the proposed method.

Read the paper · More papers on PaperTik