Optimizing fitness exercise recognition using novel neural-driven feature extraction architectures

T.S. Arthi, Sivakumar Premkumar, Ayan Sheikh, Aryan Bhatt, Avinash Tiwari · 2025

Automated recognition of fitness exercises plays a pivotal role in enhancing modern fitness applications by enabling precise activity tracking and real-time feedback. This study investigates and compares two cutting-edge techniques—YOLOv8 integrated with a Convolutional Neural Network (CNN) and MediaPipe paired with a CNN—for recognizing exercises such as Bicep Curls, Front Raises, and Shoulder Presses. A custom dataset comprising 817 annotated videos was utilized, with bounding boxes (for YOLOv8) and skeletal keypoints (for MediaPipe) meticulously labeled. The YOLOv8+CNN approach employs object detection for bounding box identification alongside CNN-based classification, incorporating convolutional layers, ReLU activations, pooling mechanisms, and dropout for enhanced prediction reliability. The MediaPipe+CNN proposes normalizing 33 skeletal keypoints per frame, after which they will be transformed back into feature vectors going through classification with CNNs to allow for effective feature extraction and to aggregate temporal features. Performance metrics comprise accuracy, precision, recall, F1-score, along with inference time. Results are that the YOLOv8+CNN pipeline was very strong in accuracy, with a high value of 92.5%, and the precision at 93.1%, while MediaPipe+CNN had faster inferences with 30ms per frame, which is way better for real-time performance. It does so by pointing out some of the trade-offs in computational efficiency versus recognition accuracy, developing scalable fitness-tracking systems, and adapting such methods to multi-user and real-time usage.

Read the paper · More papers on PaperTik