EfficientViT for Video Action Recognition
Derek Jin, Shengyou Zeng · 2024
Video action recognition is a critical challenge in a wide range of practical applications. Most of the video recognition models are computationally expensive and energy intensive. EfficientViT is a new family of vision transformers that offer high efficiency while maintaining state of the art accuracy on vision tasks such as image classification and semantic segmentation including on daily devices and cloud computing. However, the effectiveness has only been validated on images so far. In this project, we extend EfficientViT to video processing by replacing 2D convolution with 3D convolution. The updated EfficientViT model B1-r288 was successfully trained with the Epic-Kitchens-100 video dataset, achieving top 1% verb accuracy of 27.7%, a 5.4% improvement than the original model.