Semantically Plausible and Diverse 3D Human Motion Prediction
Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann, Lars Petersson, Stephen Jay Gould · arXiv (Cornell University) · 2019
We tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular approach consists of using Conditional Variational Autoencoders (CVAEs). Existing approaches to generate diverse motions either fail to capture the diversity in human motion or fail to generate diverse but semantically plausible continuations of an observed motion. In this paper, we address both of these problems by developing a new variational framework that accounts for both diversity and semantic of the generated future motion. Unlike existing approaches that sampling the latent variable is independent of the conditioning, our approach conditions the sampling process of the latent variable on the representation of the past observation, thus encouraging it to carry relevant information. Our experiments demonstrate that our approach not only yields motions of higher quality while retaining diversity, but also generates motions that preserve semantic information contained in the observed 3D pose sequence.