An Object-Oriented Deep Learning Method for Video Frame Prediction
Saad Mokssit, Daniel Bonilla Licea, Bassma Guermah, Mounir Ghogho · 2024
Predicting the next frames in a video sequence is a fundamental task in computer vision and video analysis. A great deal of success has been achieved in this task thanks to recent advances in deep learning. However, most existing techniques are tested in constrained settings such as static camera, static scene, or few moving elements. In this paper, we develop a new method of predicting the next video frame based on the estimation of transformation parameters for each object in the scene. Our method disentangles camera intrinsic motion from objects motion. First, it estimates the projection parameters for each object within the scene to align with its view in the next frame. Then, each object’s motion is captured by estimating an affine transform that aligns the warped view of the object with the ground truth next frame. The sequence of projective and affine transforms is then fed to a deep neural network (a transformer network) to predict next frames. Experiments using the Indian Driving Dataset (IDD) demonstrate the merits of the proposed method.