Action Recognition by Fusing Spatial-Temporal Appearance and the Local Distribution of Interest Points
Mengmeng Lu, Liang Zhang · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2014
The traditional Bag of Words (BOW) algorithm considers the frequency of visual words only, whereas it ignores their spatial and temporal correlations.Many methods have been designed to remedy this defect .In this paper, we propose a new descriptor to describe the local spatio-temporal distribution information of each point.This new descriptor, combined with HOG3D, is used to describe human actions.K-means clustering algorithm is introduced to generate codebook of visual words, achieving the integration of two features under the BOW model.Finally, Support Vector Machine (SVM) is used for action recognition.We extensively test our method on the standard Weizmann and KTH action datasets.The results show its validity and good performance.