Dynamic Temporal Resolution Transformers
Wei Wang, Lianfa Zhang · 2024
Action recognition is an important part of video understanding tasks, aiming to identify and classify various human actions. According to the start and end time, actions can be divided into short-term events and long-term events. Actions of different durations require dynamic time resolution adjustments based on motion complexity, which leads to fluctuations in the results obtained under the same model. We propose a DTRT (Dynamic Temporal Resolution Transformers) model to solve action recognition of different complexities. Specifically, we designed multiple temporal resolution standards, consisting of multiple independent TFM (Temporal Flash Memory) block implementation, each TFM block has the ability to dynamically extract video temporal information, and these features will be fused with the input of the next adjacent block. Therefore, the model ultimately models spatio-temporal features with information at multiple temporal resolutions. In the experimental stage, we conducted detailed ablation Studies and verified that the features modeled by DTRT under the dynamic time resolution setting can effectively represent actions of different durations to adapt to situations from fine-grained motion to long-term events. It can achieve optimal performance in the recognition of long-term events, and its overall performance is not inferior to conventional video recognition models.