Combing RGB and Depth Map Features for human activity recognition

Yang Zhao, Zicheng Liu, Lu Yang, Hong Cheng · Asia-Pacific Signal and Information Processing Association Annual Summit and Conference · 2012

We study the problem of human activity recognition from RGB-D sensors when the skeletons are not available. The skeleton tracking in Kinect SDK works well when the human subject is facing the camera and there are no occlusions. In surveillance or senior home monitoring scenarios, the camera is usually mounted higher than human subjects and there may be occlusions. Consequently, the skeleton tracking does not work well. In RGB image based activity recognition, a popular approach that can handle cluttered background and partial occlusions is the interest point based approach. When both RGB and depth channels are available, one can still use the interest point based approach. But there are questions on whether we should extract interest points independently on each channel or extract interest points from one of the channels. The goal of this paper is to compare the performances of different ways of extracting interest points. In addition, we have developed a depth map based descriptor. We show that the best performance is achieved when we extract interest points solely from RGB channels, and combine the RGB based descriptors and depth map based descriptors.

Read the paper · More papers on PaperTik