CTC Network with Statistical Language Modeling for Action Sequence Recognition in Videos
Mengxi Lin, Nakamasa Inoue, Koichi Shinoda · 2017
We propose a method for recognizing an action sequence in which several actions are concatenated and their boundaries are not given. The proposed method combines Connectionist Temporal Classification (CTC) and a statistical language model. CTC can learn the nature of each element action given no boundary information in an end-to-end manner. The statistical language model can learn the relationship between actions. We evaluate our method on the Breakfast dataset. When we use the trigram as the language model, its accuracy rate is 43.4%, which is better than the state-of-the-art ECTC method by 6.7 percentage points.