PredRNNv3: A CNN-transformer collaborative recurrent neural network for spatiotemporal prediction

Guodong Guo, You Zhou · Neurocomputing · 2025

Thanks to the development of deep learning technologies, spatiotemporal prediction learning has gained widespread attention for its great potential in various industries. Integrating CNN and RNN to capture spatiotemporal dependencies is a common video prediction strategy. Some studies have also achieved excellent results by incorporating the Transformer framework into RNN. However, current mainstream research hasn’t fully combined the advantages of CNN and Transformer to enhance the model’s ability to capture multi-scale spatiotemporal relationships. This may lead to severe distortion of the predicted results in overall or local details. Based on an in-depth analysis of the previous PredRNN model, we have developed a new simplified recurrent neural network model, termed PredRNNv3. This model employs a dual-branch architecture combining CNN and ViT to accurately capture local and global spatial dependencies. Through an ingeniously designed modulation module that integrates local and global spatiotemporal memory streams, we effectively facilitate the complementary feature learning between CNN and ViT. Without resorting to any tricks or additional loss functions, the experimental results show that the method is effective when compared to many state-of-the-art approaches.

Read the paper · More papers on PaperTik