Sports Event Video Sequence Action Recognition Based on LDCT Networks and MTSM Networks

Chunhua Xue, Jianjian Lin · IEEE Access · 2025

Due to the complexity of sports events scenes and the variability of dynamic features, the overall effect of sports video action recognition is poor. To solve this problem, the study proposes a sports video action recognition technique based on local difference feature and channel temporal. The study uses local frame extraction mode to extract spatial features and channel temporal excitation mode to extract local motion features. Considering the difficulty of hands-on recognition in complex race scenes, a motion recognition model combining 3D convolutional networks is introduced, and the extraction of motion features is realized by motion enhancement module and spatiotemporal feature aggregation module. In the LDCT network performance test, the research model outperformed similar models with 98.2% accuracy in Top-5 and 0.315 training loss in the home-made dataset. In addition, in the more complex competition scenario test, the average recognition accuracy of the research model in fencing, volleyball, equestrian, and basketball was 95.6%, which outperformed similar models. In the comparison of model recognition time consumed, the average time consumed by the research model was 0.34sm, which resulted in the shortest time consumed for sport recognition. Finally, in the comprehensive performance comparison, the research model outperformed in terms of accuracy and resource overhead. It indicated that the research technique satisfied the requirements of video action recognition for competitions. The research will provide technical support for the recognition of sports and the improvement of related techniques.

Read the paper · More papers on PaperTik