Reduced‐gate convolutional long short‐term memory using predictive coding for spatiotemporal prediction

Nelly Elsayed, Anthony S. Maida, Magdy Bayoumi · Computational Intelligence · 2020

Abstract Spatiotemporal sequence prediction is an important problem in deep learning. We study next‐frame(s) video prediction using a deep‐learning‐based predictive coding framework that uses convolutional LSTM (convLSTM) modules. We introduce a novel rgcLSTM architecture that requires a significantly lower parameter budget than a comparable convLSTM. By using a single multifunction gate, our reduced‐gate model achieves equal or better next‐frame(s) prediction accuracy than the original convolutional LSTM while using a smaller parameter budget, thereby reducing training time and memory requirements. We tested our reduced gate modules within a predictive coding architecture on the moving MNIST and KITTI datasets. We found that our reduced‐gate model has a significant reduction of approximately 40% of the total number of training parameters and a 25% reduction in elapsed training time in comparison with the standard convolutional LSTM model. The performance accuracy of the new model was also improved. This makes our model more attractive for hardware implementation, especially on small devices. We also explored a space of 20 different gated architectures to get insight into how our rgcLSTM fits into that space.

Read the paper · More papers on PaperTik