Joint Human Detection and Head Pose Estimation via Multistream Networks for RGB-D Videos

Guyue Zhang, Jun Liu, Hengduo Li, Yan Qiu Chen, Larry Steven Davis · IEEE Signal Processing Letters · 2017

We propose a multistream multitask deep network for joint human detection and head pose estimation in RGB-D videos. To achieve high accuracy, we jointly utilize appearance, shape, and motion information as inputs. Based on the depth information, we generate scale invariant proposals, which are then fed into a novel contextual region of interest pooling (CRP) layer in our deep network. This CRP has two branches to deal with contextual information for each subject. The proposed method outperforms state-of-the-art approaches on three public datasets.

Read the paper · More papers on PaperTik