ViTYoga: Vision Transformer for Real-Time Yoga Pose Estimation and Analysis

Ramesh S. Wadawadagi, A. G., Vinod L. Desai, Emmanuelle Varon, Shantakumar B. Patıl · 2025

Recognizing and analyzing yoga poses is a critical aspect of yoga practices and is vital when it is supported by technology. The importance of yoga arises from its ability to promote mental and physical well-being, stress reduction, im-proved focus, emotional stability, resilience, spiritual growth, and a life-centered perspective. As the popularity of yoga continues to grow, technological advancements are becoming increasingly vital to assist yoga instructors and practitioners in monitoring and enhancing their practice. This work elucidates the significance of automated yoga pose detection and its potential benefits for enhancing self-correction, providing real-time feedback, and maximizing yoga benefits. The proposed work utilizes Vision Transformers (ViT) for real-time detection and monitoring of yoga poses, referred to as ViTYoga. The proposed ViTYoga employs ViT model, which is pre-trained on the COCO-pose dataset and further fine-tuned on Yoga Pose Image datasets that comprise specific yoga postures. The performance assessment of proposed ViTYoga is compared and contrasted with other baseline techniques. The experimental results reveal that ViTYoga outperforms several SOTA techniques. Moreover, the proposed design offers a lightweight model that can be deployed as a mobile application and used by yoga practitioners to perform yoga at their convenience.

Read the paper · More papers on PaperTik