Online Action Detection via Temporal Dependency Modeling with Semantics Embedding
Sensen Wang, Yuehu Liu · 2024
Online Action Detection aims to identify ongoing incomplete actions from online video streams. The key is to model the inter-frame relationships related to the current action, i.e., temporal dependency. However, frames with similar appearances may contain different action semantics. Therefore, modeling the attention interaction between frame features without considering the action semantics may result in incorrect dependency between frames with similar appearances but different semantics. To address this problem, we propose to utilize action semantics and frame features to jointly model temporal dependency. Specifically, to mine frame-level action semantics, our proposed Semantic Correlation Transformer (SecTR) first extracts representative features from redundant and diverse action frames features by performing K-means clustering for each action category. Considering that a single frame only contains local action information, SecTR reconstructs the frame feature based on clustering features from the perspective of various categories, which is used to mine the action semantics of frames. Then, according to consistency with the semantics of the current frame, SecTR further embeds the frame semantics into the frame features to jointly model temporal dependency. Moreover, to further optimize the modeling quality, we propose an auxiliary task called current action anticipation, which predicts the current action based on historical frames. The effectiveness and superiority of SecTR are validated on THUMOS'14, TVSeries and HDD.