Hierarchical Hourglass Convolutional Network for Efficient Video Classification
Yi Tan, Yanbin Hao, Hao Ran Zhang, Shuo Wang, Xiangnan He · Proceedings of the 30th ACM International Conference on Multimedia · 2022
Videos naturally contain dynamic variation over the temporal axis, which will result in the same visual clues (e.g., semantics, objects) changing their scale, position, and perspective patterns between adjacent frames. A primary trend in video CNN is adopting spatial-2D convolution for spatial semantics and temporal-1D convolution for temporal dynamics. Though the direction achieves a favorable balance between efficiency and efficacy, it suffers from misalignment of visual clues with large displacements. Particularly, rigid temporal convolution would fail to capture correct motions when a specific target moves out of the reception field of temporal convolution between adjacent frames.