Mobile Action Modeling: Efficient and Effective Scenery Classification With Ball Applications
Dacheng Ma · IEEE Access · 2025
Human action recognition in ball sports presents significant challenges in computer vision. In this article, a new pipeline is designed utilizing mobile depth cameras, such as the iPhone’s LiDAR, to improve action recognition accuracy by integrating depth data into gesture analysis. Our methodology encodes human joint information captured by depth cameras, employing an intelligent accumulation of joint point data to classify various actions. The proposed pipeline offers two flexible techniques: shallow and deep joint point accumulation. Shallow accumulation prioritizes rapid detection for time-sensitive applications, while deep accumulation ensures higher accuracy for detailed 3D action reconstruction. By leveraging depth data to model regions of human activity, our system effectively distinguishes between multiple ball sport actions. It further calculates joint angles and repairs occluded points, enhancing recognition reliability. Experiments indicate the robustness of the proposed pipeline in complex scenarios. However, challenges remain, particularly in joint point detection within intricate backgrounds and under varying environmental conditions. These challenges underscore the importance of improving foreground extraction and tracking mechanisms. This work demonstrates the potential of integrating depth information with intelligent joint analysis to overcome traditional limitations in human action recognition, advancing its application in ball sports.