A Novel Video Understanding Network Based on Poolformer and Transformer

Shiqiang Shu, Hewei Yu, Jingxi Yu · 2023

This paper introduces a new video understanding network referred to as 'ViViP' (Video Vision Poolformer), which leverages poolformer and transformer techniques. It begins by encoding video frames using sequential poolformer layers. The ensuing output token is then processed by a Transformer, and the class token is subsequently used in a Multilayer Perceptron (MLP) for classification. ViViP shows notable results across multiple action recognition benchmarks, achieving the highest reported accuracy on UCF-101 and substantial accuracy on HMDB-51. Additionally, our model exhibits faster training speeds than other popular networks, such as SlowFast, TimeSformer, and TSN, allowing its usage with longer video clips. Our experiment, which compares different token-mixer modules, finds that the poolformer achieves the best classification accuracy among the examined design choices, reaching 98.4% on UCF-101 and 71.6% on HMDB-51.

Read the paper · More papers on PaperTik