Motion Guided Feature-Augmented Network for Action Recognition

Zhenxing Zheng, Gaoyun An, Qiuqi Ruan · 2020

Motion information is a crucial factor to identify human action recognition in videos. Existing state-of-the-art methods use traditional optical flow features representing the short-term motion information as the supplementary to appearance features, achieving performance improvements. However, conventional optical flow needs high computation cost and can't be optimized in an end-to-end fashion. To reduce the computation complexity and make the computation fully-differential, we utilize a representation flow layer to mimic the process of iterative flow optimization. Besides, to attach high importance to the most critical elements, motion information is used to locate motion salient objects. Toward this end, we propose a motion-augmented module to adaptively aggregate multiple optical flow features, deriving motion saliency maps that can alleviate the unreliable motion problem. Then, these motion saliency maps attend appearance features, yielding motion-augmented features. Further, joining the motion-augmented module and flow layer, we propose a Motion Guided Feature-Augmented Network to learn motion-enhanced features. Experiments on two challenging action recognition benchmarks verify that the motion-augmented module can locate salient objects that are relevant to actions, and the proposed network can achieve the competitive performance with state-of-the-art methods.

Read the paper · More papers on PaperTik