Action Recognition Based on Features Fusion and 3D Convolutional Neural Networks

Lulu Liu, Fangyu Hu, Jiahui Zhao · 2016

The paper proposes a method for human action recognition which focuses on solving the problems resulting from complex hand-crafted features. The method aggregates both spatial and temporal features and can be divided into two parts: multi-channel feature fusion and action classification. For adding motion and shape information, it firstly combines gray, optical flow and Difference of Gaussian(Dog) feature three channels for every frame in a clip. Then put the multi-channel clips into a 3D Convolutional Neural Networks(CNNs) adjusted by testing several times. The method avoids expending engineers' vigor in extracting special features from video in different scene, and experiments on the KTH action database show that it performs some better comparing with conventional algorithm and some CNNs-based methods.

Read the paper · More papers on PaperTik