Towards a generative approach to activity recognition and segmentation.

Hilde Kuehne, T. Serre · arXiv (Cornell University) · 2015

As research on action recognition matures, the focus is gradually shifting away from categorizing manually-segmented clips into basic action units to parsing and understanding long action sequences that make up human daily activities. There is a long history of applying structured models for the analysis of temporal sequences. Yet, while they seem like an obvious choice for the recognition of human activities, they have not quite reached a level of maturity in vision which is comparable to that speech recognition. With the widespread availability of large video datasets, combined with recent progress in the development of compact feature representations, the time seems ripe to revisit structured generative approaches. We propose an end-to-end generative framework which uses reduced Fisher Vectors (FVs) in conjunction with structured temporal models for the segmentation and recognition of video sequences. It shows that the overall generative properties of FVs make them especially suitable for a combination with generative models like Gaussian Mixtures. The proposed approach is extensively evaluated on a variety of action recognition datasets ranging from human cooking activities to animal behavioral analysis. It shows that the architecture, despite its simplicity, it is able to outperform complex state-of-the-art approaches on all larger datasets.

Read the paper · More papers on PaperTik