Real-Time Human Fall Detection From Video Using YOLOv11 With Pose Estimation: A Paradigm Shift Toward Efficient Transformer-Based Architectures

Adel BenAbdennour, Mahmoud Sameh, Bilal A. Khawaja, Arshad Karimbu Vallappil, Abdulmajeed M. Alenezi, Qammer H. Abbasi, Sameer Qazi · IEEE Access · 2026

The increasing global aging population and the associated risk of falls necessitate the development of efficient and reliable automated human fall detection systems. This paper presents a novel two-stage, lightweight deep learning pipeline designed for real-time fall detection. The proposed architecture combines a state-of-the-art pose estimation model, You-Only-Look-Once (YOLO)v11-Pose, with a compact transformer network for temporal classification. The methodology includes a robust data preprocessing pipeline, careful data partitioning, and systematic ablation studies across architectural components, hyperparameters, and input features. Evaluation on a combined dataset of 6,988 videos demonstrates the model’s high performance. The best F1 model achieved an accuracy of 97.96%, an F1-score of 97.74%, a Precision of 97.74%, and a Recall of 97.74% on the held-out test set. Cross-dataset validation on the LE2I dataset achieves 91.65% F1-score without fine-tuning, demonstrating generalization capability. The full pipeline requires only 142MB VRAMand achieves 76.2 FPS on a Tesla P100, exceeding real-time requirements by 2.5×, confirming its suitability for deployment in real-time monitoring systems. The findings demonstrate the effectiveness and computational efficiency of the transformer-based approach for safety-critical fall detection applications.

Read the paper · More papers on PaperTik