STNAM: A Spatio-Temporal Non-autoregressive Model for Video Prediction

Yuan Yu, Zhaohui Meng · 2024

Video prediction requires efficient models capable of forecasting future frames which is a crucial task in various domains. However, many current methodologies are based on autoregressive mechanism, suffering from low computing efficiency, error propagation and difficulty in parallel processing of data. With an emphasis on efficiency, we propose the Spatio-Temporal Non-autoregressive Model (STNAM) designed for video prediction tasks. This model aims to achieve superior computational efficiency and reduced error accumulation compared to conventional methodologies. The STNAM is grounded in encoder-prediction-decoder framework with a Spatio-Temporal Attention and a Positional encoding. Experimental evaluations on benchmark video datasets showcase the efficacy of the proposed model. It demonstrates competitive performance in predicting video sequences, establishing its potential for real-time video forecasting applications.

Read the paper · More papers on PaperTik