Selecting active frames for action recognition with vote fusion method
Hoang Tieu Binh, Ma Thi Chau, Akihiro Sugimoto, Bui The Duy · 2018
Recent applications of Convolutional Neural Networks, especially 3-Dimensional Convolutional Neural Networks (3DCNNs) for human action recognition (HAR) in videos have widely used. In this paper, we use a multi-stream framework which is a combination of separated networks with different kinds of input generated from a unique video dataset. We study various methods for extracting best frames in videos for action representation and find the way to fuse multiple networks for better recognition. We make the following work: first, we propose a method to extract the active frames (called Selected Active Frames - SAF) from videos to build datasets for 3DCNNs in video classification problem. Secondly, we propose a mixed fusing approach called Vote fusion which is considered as an effective fusion method for ensembling multi-stream networks. We evaluate the proposed approach to solving action recognition problem. We carry out this approach on three well-known datasets (KTH, HMDB51, and UCF101). The results are also compared to the state-of-the-art results to illustrate the efficiency and effectiveness in our approach.