TDP: Temporal Dynamic Pooling — A New Method for Temporal Action Localization
Lei Li, Lihong Ma, Jing Tian · 2018
Temporal action localization is a very important yet challenging task, since video in real applications is usually long, unconstrained and untrimmed, containing multiple action instances plus video content of complex background scenes. The task recognizes not only the action class but also the start frame and the end frame of each action instance. To address this issue, we propose a novel temporal dynamic pooling (TDP) network to determine the temporal boundaries of an action instance and to select key frames of a video through a sequential prediction of importance score, which is obtained by score prediction network realized based on residual learning. By pooling only a few key frames containing discriminative information, we could obtain the salient pooled vector for action recognition. We conduct experiments on a challenging temporal action localization dataset, THUMOS 2014. When setting Intersection-over-Union (IoU) threshold to 0.5 in evaluation, we significantly achieve superior performances compared with the most state-of-the-art methods by increasing mAP from 22.7% to 29.8%. Meanwhile, our TDP network demonstrates a very high efficiency with the ability to process 550 fps on a single CPU.