Personalized pose estimation for body language understanding

Zhengyuan Yang, Jiebo Luo · 2017

To achieve high accuracy and stability in human pose estimation from videos, we propose a personalized model with a specially designed ConvNet structure and a visual similarity based iteration step. This model consists of: 1) a fully convolutional network with spatial fusion architecture to boost the accuracy of single-frame based joint predictions, 2) optical flow-based refinement to incorporate motion and temporal information, and 3) iterative personalized annotation to boost the reliability of the joint predictions. For benchmarking, our model outperforms the state-of-the-art on the public pose estimation datasets Chalearn and FLIC. Moreover, our model performs the best on a new psychiatric conversation dataset for computer vision based body language and emotion study.

Read the paper · More papers on PaperTik