Vision-Based Human Motion Analysis

Angela Yao · Repository for Publications and Research Data (ETH Zurich) · 2012

Interpreting human activity from video is at the core of a wide spectrum of applications such as content-based indexing, intelligent surveillance, human-computer interfacing and sports video analysis. Cheap hardware and growing storage capacity has led to an explo-sion of video data and there is a critical need for machine vision algorithms that automat-ically analyze video content. This thesis provides a collection of methods for video-based human action recognition, i.e. the application of semantic labels to a person’s movements over time in a video se-quence. We present two approaches for this task, one appearance-based and one pose-based. The appearance-based method uses no structural modeling of the human body and relies only on the statistical distribution of appearance features such as edges, shapes and flow to classify actions. The pose-based method, on the other hand, explicitly estimates a 3D articulated pose of the body and classifies the action based on geometric relations between specific joints in a single pose or a short sequence of poses. In addition to determining action labels, we examine how action recognition could be leveraged to help with the closely related task of human pose estimation. We integrate ac-tion recognition and pose estimation into a single system, taking output from appearance-based action recognition as a prior for 3D pose estimation. The estimated poses are then used to for pose-based action recognition to refine the action label. Finally, we examine the temporal aspect of labeling actions and propose a method to both segment and classify actions from a continuous stream of body poses.

Read the paper · More papers on PaperTik