Interaction Aware Relational Representations for Video Prediction
Rei Tamaru, Prarinya Siritanawan, Kazunori Kotani · 2021 IEEE International Conference on Systems, Man, and Cybernetics (SMC) · 2021
Video prediction is an active machine learning problem to use past information in a video sequence to acquire human-like understanding and then predicting future consequences of object states and actions. The existing prediction frameworks integrated the decomposition and disentanglement techniques to observe object interaction and use them to predict future video scenes. However, the previous works did not consider physical interaction among objects in the prediction. Thus, this research utilizes the physical reasoning concept to represent object dynamics in the real world and estimate future sequences enclosing object dependencies. This paper addresses the investigation of object interaction using stochastic video prediction with physical reasoning representation. We propose a self-supervised framework called Relational Prediction Auto-Encoder (RPAE). Extensive experiments demonstrate that the proposed RPAE can effectively improve the generation and prediction of the near future sequences. We also confirmed the predicted object dynamics by measuring the velocity of each object and its physical interaction in the experiments.