Complex Event Detection Using Joint Max Margin and Semantic Features

Iman Abbasnejad, Sridha Sridharan, Simon Denman, Clinton Fookes, Simon Lucey · 2016

In this paper the problem of complex event detection is addressed. Existing event detection methods are limited to features that are extracted from local spatial or spatio-temporal patches from the videos. However, this makes the model more vulnerable to the events that have similar concepts with different actions e.g. "Open drawer" and "Open cupboard". Furthermore, current methods typically assume that events have already been segmented from the video stream, and do not generalize well to events with unknown starting and ending locations. In this work, in order to address the aforementioned limitations we present a novel model based on the combination of semantic and temporal features extracted from video frames. We train a max-margin classifier on top of the extracted features in an adaptive framework that is able to detect the events with unknown starting and ending locations. Our model is based on the Bidirectional Region Neural Network and large margin Structural Output SVM. The generality of our model allows it to be simply applied on different labeled and unlabeled datasets. We finally test our algorithm on three challenging datasets: "UCF 101-Action Recognition", "MPII Cooking Activities" and "Hollywood"; and we report the state-of-the-art performance.

Read the paper · More papers on PaperTik