Towards temporal adaptive representation for video action recognition
Junjie Cai, Jie Yu, Francisco Hideki Imai, Qi Tian · 2016
Action recognition has been one of the challenging problems in the computer vision community. Most of the recent research work in this area exploits the motion features captured by dense trajectory descriptors. On the other hand, static image classification has seen the rise of deep learning architectures, with evidence that the output of intermediate layers could be successfully employed as a low level descriptor for new learning tasks. However, the same level of classification success has not yet translated to the video domain. In this paper, we investigate to jointly combine dynamic trajectory features and static deep features that enhance the distinctiveness of the classifiers. We also propose a Temporal-Pyramid-Pooling strategy with intermediate layer deep features for improving action classification performance. Extensive empirical evaluations are provided to corroborate the effectiveness of the proposed framework on real-world untrimmed video datasets.