On Understanding of Spatiotemporal Prediction Model

Xu Huang, Xutao Li, Yunming Ye, Shanshan Feng, Chuyao Luo, Bowen Zhang · IEEE Transactions on Circuits and Systems for Video Technology · 2022

Recently, explainable artificial intelligence has received considerable attention. Most existing studies are focusing on the tasks of CNNs-based image classification and RNNs-based time series analysis. In this paper, we pay attention to the more complicated spatiotemporal predictive learning task (SPLT), where both the spatial and temporal information play important roles. To explain the internal mechanism of spatiotemporal prediction models, we propose a comprehensive analysis method. Specifically, with a typical encoder-decoder framework, we focus on two core issues of SPLT: image generation and spatiotemporal dynamics. For the first issue, we develop aquantitative channel perturbationmethod to explore the importance of features to prediction. Furthermore, we propose a technique called thesynthesis of multiple independent componentsto analyze how these features generate the prediction. According to the experimental results, thecoarse- and fine-grainedsynthesis (CFGS) mechanism is drawn for image generation in SPLT. For the second issue, we propose astate decompositiontechnique and astate expansiontechnique to disentangle coupled signals in the spatiotemporal dynamical system. This helps us to explore the mechanism of forming motion. Moreover, to diagnose the movement of a particular region during analysis, we propose a fluorescent stamp-based technique. By observing extensive experimental results, we summarize a collaboration mechanism to explain how the motion is formed in SPLT, namely, theextending the present and erasing the past (EPEP)mechanism. To the best of our knowledge, this is the first work to interpret the internal mechanism of SPLT models.

Read the paper · More papers on PaperTik