Classification and Temporal Localization for Human-Human Interactions
Ngoc Nguyen, Atsuo Yoshitaka · 2016
Recognition of human-human interactions is one of the most important topics since it has great scientific importance and many potential practical applications such as surveillance, and automatic video indexing. Previous approaches have only concentrated on classification and put less effort into localization of human interactions. In addition, they rely on hand-designed features (e.g. SIFT, HOG), or human poses or human joints to model human interactions. A disadvantage of such approaches is that it is difficult and time consuming to extend these features to different datasets in the real world. In this paper, we approach the problem of human interaction classification and temporal localization with unsupervised feature learning. Motivated by the well-known Independent Subspace Analysis (ISA) in natural image statistics and the convolution technique, we introduce a three-layer convolutional ISA network to learn hierarchical invariant features from videos. Using the invariant features learned by the three-layer convolutional ISA network, we build a bag-of-features (BOF) representation for videos. We then apply Support Vector Machine (SVM) to classify human interactions, and employ a sliding window technique to localize interactions temporally. We also evaluate the performance of classification and temporal localization on video sequences of the UT-Interaction dataset and Hollywood dataset. The encouraging results on classification show that our three-layer convolutional ISA network is able to learn features which are effective to represent complex activities such as human interactions in realistic environments. Although temporal localization results are insufficient for real applications, it is a first step for further research in localization of human interactions.