Scaled Weight-Sharing Convolutional Network for Memory-Efficient Video Frame Prediction

Jiaqian Xie, Shun-ichi Sekiguchi, Wataru Kameyama · 2024

Recent developments in pure convolutional neural network (CNN) architectures have significantly advanced spatiotemporal modeling for video frame prediction. However, traditional architectures like PredNet require extensive parameter tuning and long training times. The SimVP and its variant models have addressed these issues by using purely CNN-based architectures, reducing computational complexity and improving training efficiency. Inspired by their works, in this paper, we enhance CNN-based video frame predictors by scaling dimensionality and the number of convolutional layers to capture more complex relationships between frames. To maintain efficiency, we introduce a weight-sharing scheme that reduces model parameters while maintaining high prediction quality. Extensive experiments on three public datasets demonstrate our approach's efficacy and efficiency, achieving significant improvements in frame prediction accuracy and computational performance while optimizing resource usage.

Read the paper · More papers on PaperTik