Improving multi-view human action recognition with spatial-temporal pooling and view shifting techniques

Tuan-Dung Le, Thi-Oanh Nguyen, Thanh-Hai Tran · 2017

This paper presents a solution to improve performance of human action recognition from multiple camera views. For each camera view, we started by investigating a bag of words model that consists of STIP features to capture motion, a random forest for feature quantization and a SVM for action classification as baseline. However, to avoid background effect, we take STIP features only in the moving regions detected by background subtraction technique. Then, as some actions are very similar (interclass similarity), they discriminate against each others by some minor motions of body part (hand or foot) or/and by the order of movement during the action, we adopt a spatial-temporal pooling strategy of STIP features to take this difference into account. Finally we propose a strategy of shifting views in testing phase to deal with difference of camera viewpoints from training phase. The result from each view will be combined by late fusion. The proposed method has been evaluated on a benchmark multi-view dataset WVU. Experimental results show that this method achieves 92.28% of accuracy, that outperforms the baseline by 23.92%. It is evenly better than a method using advanced convolution neural network by 2.28%.

Read the paper · More papers on PaperTik