Generating Future Frames with Mask-Guided Prediction
Qian Wu, Xiongtao Chen, Zhongyi Huang, Wenmin Wang · 2020
Current approaches in video prediction tend to hallucinate the future frames directly or learn global motion transformation from the entire scene. However, it is difficult for these methods without instance-aware mechanism to learn the underlying structures, dynamics and appearances of foreground and background elements simultaneously, especially when it comes to long-term prediction. In this paper, we propose an explicit instance-level prediction approach to tackle this issue and present a novel mask-guided dual network. We utilize instance masks to extract active objects from the videos, and design two LSTM branches to predict the future dynamics and appearances for objects and backgrounds individually. Superior than most recent skeleton-aided methods that only focus on single human object with two-stage procedure, our proposed network can predict instances from other categories and be trained end-to-end with a joint loss. We evaluate our approach on KTH, Penn Action and Running Horse datasets, and achieve promising results in both quality and quantity.