A winter sports video classification method and system based on 3D-video Swin transformer
Xiaolin Yuan · 2023
Videos of winter sports can be a very good resource for learning about sports in China. In this paper, we address the issues of a small winter sports video dataset and low classification accuracy by proposing a fast video classification approach that combines a transformer and a convolutional neural network. The 3D-Video Swin Transformer model is built using the resnet3D feature extraction component and the Video Swin Transformer, and it improves local and global modeling capabilities through a multi-headed self-attention mechanism. Convolutional operations are used at the network front-end to make for the Transformer's lack of inductive bias, enhancing the network's local modeling capabilities. Convolution operations are utilized at the network's front end to make for Transformer's lack of inductive bias, hence increasing the network's capacity for local modeling and decreasing the model's dependency on vast amounts of data. The experimental results indicate that the 3D-Video Swin Transformer model may achieve an accuracy of up to 76.43 percent on the winter sports video dataset developed in this paper. The classification impact is also substantially stronger, and this accuracy is also 1.15 percent higher than that of the Video Swin Transformer Network. Additionally, we develop and implement a winter sports video classification system based on the Milvus database to facilitate user interaction and enable the submission, categorization, and recommendation of winter sports movies.