Enhancing Human Pose Estimation: A Data-Driven Approach with MediaPipe BlazePose and Feature Engineering Analysing
Vishakha Arya, Srabanti Maji · 2024
Computer vision has seen a surge in interest and research on human activity recognition using pose estimation techniques. It aims to classify the poses of human body joints via images and videos. Human Pose estimations work on 2D and 3D skeleton-based action recognition of human body landmarks, tailored explicitly to real-time on-device inferences. The study investigated the core components of the MediaPipe architecture, analysing mechanisms for keypoint detection, part affinity field generation, and real-time processing. It follows a two-step ML pipeline estimator and detector for pose estimation; the estimator directly computes key points for faster inference, whereas it utilises the skeleton-based approach for detecting, leveraging the topology of 33 landmarks derived from BlazePose. First, our model implements the BlazePose lightweight pose estimation convolutional neural network on IntelliRehabDS-IRDS, a rehabilitation dataset in which Kinect motion sensors record gestures. The key points explored in this paper's specific research include (1) Identifying human activities by detecting specific vital points and assessing the accuracy of their movements, which have been empirically shown to produce reliable outcomes to help enhance overall effectiveness. (2) t-Distributed Stochastic Neighbor Embedding and Principal Component Analysis is utilised to visualise data in 2D and 3D space, preserving 96.57% of the total explained variance. Feature_importance is used for feature selection: HandTipRight_3y 0.035, HandLeft_3y 0.02, and so on.