Optimizing action segmentation with linear self-attention and pyramidal pooling
Jiamin Fu, Zhihong Chen, Haiwei Zhang, Yuxuan Gao · 2024
Continuous action segmentation is a challenging task in video semantic understanding, aims to temporally segment unedited long videos. Current state-of-the-art methods combine time-domain convolution with self-attention mechanisms to capture temporal correlations, achieving high-accuracy frame-level classification and reducing over-segmentation during prediction. However, these models rely on multiple decoding modules and complex self-attention mechanisms, which, while improving frame accuracy, incur high computational costs, particularly when dealing with long video sequences. To address this issue, we propose a novel hybrid model that optimizes the temporal convolutional model with a simplified extended linear self-attention layer. This design enables the model to focus more effectively on action changes at key moments in a video sequence while significantly reducing computational complexity. Furthermore, we introduce an attention enhancement module with a pyramidal pooling structure to improve the model's ability to capture actions at multiple scales. By leveraging these innovations, our hybrid model achieves a better balance between accuracy and computational efficiency, making it more suitable for real-world applications involving long video sequences. The experimental results demonstrate that our model achieves an accuracy of 86.0% on the 50Salads dataset and 84.6% on the GTEA dataset, outperforming existing techniques such as MS-TCN++ and ASFormer in terms of both accuracy and computational efficiency. Notably, our model requires only 0.98M parameters, a significant reduction compared to ASFormer's parameter count. These advantages make our model particularly well-suited for practical applications involving long video sequences, where computational efficiency and accuracy are crucial.